Site Reliability Engineer - Memphis
SpaceXAI - Memphis, TN
Hiring: Site Reliability Engineer - Memphis Company: SpaceXAI Location: Memphis, TN Job Posted Time: 2026-09-10 10:23:39 Target Skills & Keywords : Bash, C, C++, Java, Python, Rust About the job Experience: •Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries. •Operational familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry. •Prior work in a fast-paced startup or tech company like SpaceXAI. Required Skills: •Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure. •Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene. •Run blameless postmortems and drive corrective actions to closed, not filed. •Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries. •Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects). •Define error budgets and availability objectives at campus and service boundaries as adopted by the business. •Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus. Qualifications: •Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience). •5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments. •Proven large-scale incident command experience and calm technical leadership on a bridge. •Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality. •Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them. •Excellent problem-solving skills with a data-driven approach to reliability engineering. •Demonstrated capacity to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!