Senior Software Engineer - GPU Local AI Platforms
NVIDIA - Seattle, WA
Hiring: Senior Software Engineer - GPU Local AI Platforms Company: NVIDIA Location: Seattle, WA Job Posted Time: 2026-09-09 17:59:08 Employment Type: Full-time Target Skills & Keywords: GPU Computing, CUDA, Triton, LLM Inference, Attention Mechanisms, KV-Cache Management, Continuous Batching, Quantization, Tensor Parallelism, NCCL/RCCL, Docker/OCI, NVIDIA Container Toolkit, Python, C++, Performance Analysis, CI/CD Experience: - 12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference - Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) - Working knowledge of LLM inference internals: attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism - Experience characterizing multi-node inference behavior and collective communication primitives (NCCL/RCCL) Required Skills: - Strong Python or C++ programming, software design, and software engineering skills - GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) — understanding how thread blocks, memory hierarchy, and warp execution affect real-world performance - Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization, runtime configuration, NVIDIA Container Toolkit - Strong analytical skills for performance analysis and hardware limit mapping - Ability to track and evaluate innovations in leading open-source LLM inference frameworks - Experience developing and maintaining developer-facing inference recipes and model validation workflows Qualifications: - BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience - Proven ability to analyze how new model architectures and inference algorithms map onto GPU architecture - Experience engaging with community and partners on model bring-up questions - Strong technical communication skills as a technical point of contact for hardware-specific inference issues Compensation: Not specified in the provided job text. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!