Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud
NVIDIA - Seattle, WA
Hiring: Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud Company: NVIDIA Location: Seattle, WA Job Posted Time: 2026-09-02 17:59:08 Employment Type: Full-time Target Skills & Keywords: Kubernetes, Distributed Systems, Containers, Systems Performance & Scalability, GPU Operator, Network Operator, Node-Feature-Discovery, Topograph, DRA Driver NVIDIA GPU, NVSentinel, Grove, Gateway API Inference Extension, Confidential Containers (CoCo), DSX, AI Infrastructure, Cloud Platforms Experience: - Deep expertise in distributed systems, Kubernetes, containers, and systems performance and scalability - Broad, hands-on experience across the stack, including GPU operators, device plugins, distributed inference serving, and major cloud platforms - Proven ability to own hard technical problems at large scale and help shape how AI infrastructure runs in production - Experience leading end-to-end performance and scalability analysis across Kubernetes-based accelerated runtime stacks (control and data planes) - Experience designing and contributing upstream architectural changes to the Kubernetes control plane and related projects for hyperscale cluster sizes - Experience improving container startup and cold-start latency for low-latency inference scaling across thousands of GPU nodes - Experience advancing scalability and performance of confidential containers (CoCo) on Kubernetes - Experience using large-scale simulation infrastructure (e.g., DSX) to model full AI-factory deployments and validate scalability across thousands of simulated GPUs - Experience collaborating with AI researchers, developers, customers, and upstream communities to design automated, at-scale workload tests and build monitoring/analysis tooling Required Skills: - Kubernetes (control plane and data plane) - Distributed systems - Containers and container orchestration - Systems performance and scalability engineering - GPU Operator, Network Operator, node-feature-discovery, topograph, dra-driver-nvidia-gpu, nvsentinel - Open-source contribution (e.g., Grove, gateway-api-inference-extension) - Confidential Containers (CoCo) - Large-scale simulation and modeling (DSX) - Automated at-scale workload testing and monitoring/analysis tooling - Multi-node training/inference architecture design - Cloud platforms Qualifications: - Deep expertise in distributed systems, Kubernetes, containers, and systems performance and scalability - Broad, hands-on experience across the stack, including GPU operators, device plugins, distributed inference serving, and major cloud platforms - Strong track record of owning and solving hard technical problems at large scale - Ability to design and contribute upstream architectural changes to Kubernetes and related projects - Experience with confidential containers and encrypted inference workloads - Experience with large-scale simulation infrastructure for AI-factory deployments - Strong collaboration skills with AI researchers, developers, customers, and upstream communities Compensation: Not specified in the provided job text. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!