Senior Site Reliability Engineer - HPC

NVIDIA - Austin, TX

Hiring: Senior Site Reliability Engineer - HPC Company: NVIDIA Location: Austin, TX Job Posted Time: 2026-09-09 17:59:08 Employment Type: Full-time Target Skills & Keywords: Site Reliability Engineering (SRE), High-Performance Computing (HPC), Slurm, LSF, Kubernetes, Infrastructure as Code (IaC), CI/CD, Multi-Cloud (AWS, GCP, OCI), Observability, AIOps, Python, Go, Perl, Ruby, Capacity Management, Incident Response, Root Cause Analysis (RCA) Experience: - 5+ years of professional experience building and supporting critical services - Experience supporting large-scale HPC clusters using Slurm, LSF, or Kubernetes clusters, including setup, tuning, and troubleshooting - 5+ years of coding/scripting experience in at least two high-level programming languages such as Python, Go, Perl, or Ruby - Experience mentoring other engineers and influencing technical direction through design reviews, architecture documents, and partnership with product and leadership Required Skills: - Proficiency in modern CI/CD techniques and Infrastructure as Code (IaC) for managing services - Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability, or data-driven operations (AIOps/ML-driven signals) - Proficient in monitoring, metrics, container management, and log collection tools - Ability to own SRE solutions end-to-end, from design and implementation to operation and continuous improvement - Experience delivering solutions in globally distributed, multi-cloud hybrid environments (On-prem, AWS, GCP, OCI) - Design for failure with redundancy, failure domains, progressive delivery, and strict change control - Capacity management and planning - Detecting performance issues and recommending solutions - Collaboration with various teams in a fast-paced environment - Participation in on-call, incident reviews, root cause identification, and producing high-quality RCA reports Qualifications: - B.S. degree in Computer Science or related technical field (or equivalent experience) - Creative problem solver with excellent debugging skills - Strong communication and documentation abilities Compensation: Not specified in the provided job text. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!