Senior Software Engineer - GPU Local AI Platforms
NVIDIA - Westford, MA
Hiring: Senior Software Engineer - GPU Local AI Platforms Company: NVIDIA Location: Westford, MA Job Posted Time: 2026-09-09 17:59:08 Employment Type: Full-time Target Skills & Keywords: GPU Computing, CUDA, Triton, LLM Inference, Tensor Parallelism, KV-Cache, Quantization, NCCL, Docker, Python, C++, CI/CD, Model Validation Experience: - 12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference - Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) - Working knowledge of LLM inference internals: attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism - Experience characterizing multi-node inference behavior and collective communication primitives (NCCL/RCCL) - Track record of evaluating open-source LLM inference frameworks and model architectures Required Skills: - Strong Python or C++ programming, software design, and software engineering skills - GPU kernel development and optimization — understanding how thread blocks, memory hierarchy, and warp execution affect real-world performance - Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization, runtime configuration, NVIDIA Container Toolkit - Performance analysis mapping theoretical hardware limits (memory bandwidth, FLOP/s, interconnect throughput) to observed inference throughput, latency, and utilization - Model validation workflow ownership: architecture compatibility assessment, inference recipe development, performance characterization - Strong analytical skills Qualifications: - BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience - Ability to track and evaluate innovations in leading open-source LLM inference frameworks - Experience analyzing how new model architectures and inference algorithms (attention variants, MoE routing, speculative decoding, multi-token prediction, quantized inference) map onto GPU architecture - Ability to develop and maintain developer-facing inference recipes with automated staleness detection and CI feedback loops - Strong communication skills to engage with community and partners on model bring-up questions Compensation: Not specified in the job posting. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!