Engineering Manager, Inference Benchmarking — AI Perf
NVIDIA - Austin, TX
Hiring: Engineering Manager, Inference Benchmarking — AI Perf Company: NVIDIA Location: Austin, TX Job Posted Time: 2026-09-15 17:59:08 Employment Type: Full-time Target Skills & Keywords: LLM Inference, Benchmarking, AIPerf, vLLM, TRT-LLM, SGLang, Kubernetes, GPU Telemetry, DCGM, PyNVML, Prometheus, ZMQ, Distributed Systems, ML Tooling, Performance Infrastructure, Open-Source, Team Leadership Experience: - 8+ overall years of software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems - 3+ years of engineering leadership experience as a tech lead, TLM, or engineering manager - Proven track record of collaborating across multi-functional groups and delivering production-quality output in high-velocity, high-external-visibility environments Required Skills: - Deep understanding of LLM inference mechanics — TTFT, ITL, KV caching, Prefill/Decode, speculative decoding — and the ability to reason about measurement correctness and reproducibility - Driving the technical roadmap for AIPerf's core infrastructure: load generation, ZMQ-based microservices, GPU telemetry (DCGM/PyNVML), Prometheus metrics, statistical confidence intervals, and Kubernetes-native deployment - Advising upstream engine integrations involving vLLM, TRT-LLM, and SGLang in partnership with NVIDIA's Dynamo and NIM teams - Hiring, mentoring, and growing a team of senior engineers operating in a high-velocity open-source environment with active external contributors worldwide - Taking ownership for the accuracy and statistical soundness of benchmark results Qualifications: - Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent experience - Extensive experience with vLLM, TRT-LLM or SGLang internals along with contributions to their upstream projects (preferred) - Experience building Kubernetes-native infrastructure including operators, Helm charts, and GPU observability tooling (DCGM, dcgm-exporter, PyNVML) (preferred) - Background in competitive benchmarking frameworks such as MLPerf or equivalent industry-standard experience (preferred) Compensation: Not specified Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!