Senior Site Reliability Engineer - HPC

NVIDIA - Durham, NC

Hiring: Senior Site Reliability Engineer - HPC Company: NVIDIA Location: Durham, NC Job Posted Time: 2026-09-09 17:59:08 Employment Type: Full-time Target Skills & Keywords: Site Reliability Engineering, HPC, Slurm, LSF, Kubernetes, Infrastructure as Code, CI/CD, Multi-Cloud, AWS, GCP, OCI, Observability, AIOps, Python, Go, Perl, Ruby, Capacity Management, Incident Response Experience: - B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services - Experience supporting large-scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting - 5+ years of coding/scripting experience in at least two high-level programming languages such as Python, Go, Perl, or Ruby - Experience mentoring other engineers and influencing technical direction through design reviews, architecture documents, and strong partnership with product and leadership Required Skills: - Proficiency in modern CI/CD techniques and Infrastructure as Code (IaC) for managing services - Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention - Proficient in monitoring, metrics, container management, and log collection tools - Creative problem solver with excellent debugging skills and strong communication and documentation abilities - Ability to own SRE solutions end-to-end, from design and implementation to operation and continuous improvement - Design for failure with redundancy, failure domains, progressive delivery, and strict change control - Conduct capacity management and planning to meet ongoing operational needs - Detect performance issues and recommend solutions to maintain world-class service quality - Collaborate with various teams in a fast-paced environment to ensure seamless project completion - Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports Qualifications: - B.S. degree in Computer Science or related technical field (or equivalent experience) - 5+ years professional experience building and supporting critical services - Experience with large-scale HPC clusters (Slurm, LSF, or Kubernetes) - 5+ years of coding/scripting experience in at least two high-level programming languages - Proven ability to mentor engineers and influence technical direction - Strong communication and documentation abilities Compensation: Not specified in the job posting. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!