Senior Machine Learning Engineer - LLM Quantization & Deployment

XPENG - Santa Clara, CA

Hiring: Senior Machine Learning Engineer - LLM Quantization & Deployment Company: XPENG Location: Santa Clara, CA Job Posted Time: 2026-09-13 03:28:52 Employment Type: Full-time Target Skills & Keywords : C++, Deep Learning, LLM, Machine Learning, Make, ONNX, PyTorch, Python, R, TensorRT, TestNG, Transformers, vLLM About the job Experience: •1-3 years of industry experience. Open to new graduates. •3 years of industry experience. Open to new graduates. Required Skills: •Develop VLA inference models, ensure numerical consistency with training models, and productionize LLM quantization methods, including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques. •Develop production-quality Python code with strong testing, observability, reproducibility, and failure handling. •Build robust model export, calibration, benchmarking, validation, and deployment pipelines. •Engage early with the VLA model research team to establish performance estimates and prove model feasibility. •Curate evaluation datasets and establish a comprehensive metric suite to systematically benchmark VLA performance. •Analyze numerical errors, accuracy regressions, and performance trade-offs. •Develop PTQ and QAT orchestration workflows. •Serve as the primary interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off. Qualifications: •Master in CS/CE/EE, or equivalent, with 1-3 years of industry experience. Open to new graduates. •In-depth knowledge of Transformer architectures and LLM inference. •Hands-on experience quantizing or deploying deep learning models in production. •Proficiency with PyTorch and at least one inference or compilation stack. •Strong Python programming and software engineering skills. •Demonstrated capacity to work effectively across research, systems, infrastructure, and product teams. •Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment. •Practical experience utilizing weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization. •Practical experience utilizing AWQ, GPTQ, SmoothQuant, or related methods. •Strong numerical analysis and systems engineering skills. Compensation: •$174,720 - $295,680 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!