Staff Site Reliability Engineer-Production Operations
Rivian and Volkswagen Group Technologies - Palo Alto, CA
Hiring: Staff Site Reliability Engineer-Production Operations Company: Rivian and Volkswagen Group Technologies Location: Palo Alto, CA Job Posted Time: 2026-09-10 11:57:08 Employment Type: Full-time / Hybrid Target Skills & Keywords : Datadog, Embedded Systems, LLM, Python, Systems Engineering About the job Experience: •6+ years of experience in SRE, production/platform engineering, or systems engineering for large-scale distributed systems, including senior technical leadership or lead responsibilities. Required Skills: •Lead incident coordination and communication •Facilitate modern, blameless post-incident learning •Own and prioritize systemic action items •Manage the backlog of action items generated by reviews. While ProdOps does not write the fixes itself, you will prioritize, track, and drive these items to closure in partnership with TPMs and development teams, keeping leadership focused on customer impact. •Once the fixes ship, measure and verify that they actually prevent recurrence. Drive accountability for outcomes and feed the results back into the Novel Incident Rate. •Build the automation and observability backbone •Design and build the systems, tooling, and AI-agent workflows that automate incident triage and administrative toil. Improve observability, including instrumentation, alerting, dashboards, and SLOs, so incidents are detected faster and understood more deeply. Write and review code where it multiplies the team's impact. •Lead and grow the team Qualifications: •Proven incident command experience: you have coordinated high-severity, cross-functional incidents and led communication with executives and external stakeholders under pressure. •Strong systems engineering fundamentals, including distributed systems, networking, cloud infrastructure, and an instinct for how complex systems fail. •Hands-on coding ability (e.g., Python, Go, or similar) sufficient to build automation, tooling, and integrations. This is not a code-free management role. •Deep observability expertise: instrumentation, metrics, logging, tracing, alerting, dashboards, and SLO/SLI design (Datadog or comparable platforms). •Fluency in modern reliability and post-incident practice: blameless reviews, Learning From Incidents (LFI), HOWIE, and systemic (non-single-root-cause) analysis. •A demonstrated bias toward eliminating toil through automation, and enthusiasm for using AI agents as force multipliers. •Excellent written and verbal communication; able to translate technical detail for executive and cross-company audiences. •Exposure to Product Security or working closely with security teams during incidents. •Track record building an SRE or reliability function from an early stage. Compensation: •$186,000 - $255,750 / year •Competitive benefits and rewards package •In addition to a competitive base salary, full-time positions may be is eligible to participate in our annual company performance bonus program Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!