Site Reliability Engineer

SpaceXAI - Palo Alto, CA

Hiring: Site Reliability Engineer Company: SpaceXAI Location: Palo Alto, CA Job Posted Time: 2026-09-11 12:23:52 Target Skills & Keywords : ArgoCD, C++, CI/CD, Grafana, Kubernetes, PagerDuty, Prometheus, Pulumi, Root Cause Analysis, Rust, Terraform About the job Experience: •Operational familiarity with data center electrical, cooling, and network systems, such as liquid-cooling and high-bandwidth interconnects. Required Skills: •Maintain and improve the reliability and uptime of xAI’s on-premises and cloud-based data center environments, including high-density GPU clusters for AI training. •Design, implement, and manage monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, PagerDuty). •Develop and maintain infrastructure-as-code (Pulumi, Terraform) and continuous deployment pipelines (Buildkite, ArgoCD). •Participate in on-call rotations, respond to incidents, perform root cause analysis, and drive post-mortem processes. •Analyze system performance, forecast capacity needs, and optimize resource utilization for massive AI/ML workloads. •Partner cross-functionally with hardware, networking, and software engineering teams to design and implement resilient, scalable solutions, such as RDMA fabrics and liquid-cooling systems. •Create and maintain documentation and standard operating procedures. •Contribute to the efficiency of AI training pipelines by identifying and mitigating bottlenecks in compute, storage, and networking at unprecedented scales. Qualifications: •Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent experience). •5+ years in site reliability engineering, data center operations, or large-scale infrastructure management. •Expert-level knowledge of Kubernetes (on-prem and cloud), infrastructure-as-code tools (Pulumi, Terraform), and CI/CD systems (Buildkite, ArgoCD). •Proficiency in at least one systems programming language (Rust, C++, Go) and strong scripting/automation skills. •Comprehensive expertise in monitoring and observability technologies. •Strong troubleshooting skills across hardware, networking, and distributed software systems. •Proven experience with incident response, including on-call rotations, rapid incident resolution, root cause analysis, and implementation of preventative measures. •Excellent communication and documentation skills, with the ability to share knowledge concisely and accurately. •Certifications in SRE, Kubernetes, or data center operations. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!