Senior Staff Lead Site Reliability Engineer (R5803)

Shield AI - San Diego, CA

Hiring: Senior Staff Lead Site Reliability Engineer (R5803) Company: Shield AI Location: San Diego, CA Job Posted Time: 2026-09-10 10:33:14 Employment Type: Full-time Target Skills & Keywords : AI, AWS, Kubernetes, Python, Systems Design About the job Experience: •7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields Required Skills: •Hivemind is looking for an experienced SRE lead to help drive the establishment of our SRE function. •Establish and mature the reliability practices used across our cloud infrastructure and platform services. You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and ensure that production systems can be operated and recovered predictably. •Act as a thought-leader and mentor within the Cloud Engineering and Reliability teams to level up teammates and encourage building with a reliability-first mindset. •Define and implement SLIs, SLOs, and other measures of service reliability •Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services •Lead technical response to complex incidents and drive root-cause analysis through resolution •Identify recurring failure modes and work with engineering teams to eliminate them •Improve system resilience through automation, testing, capacity planning, and failure recovery Qualifications: •Demonstrated capacity to diagnose complex failures across applications, infrastructure, networking, and dependent services •Background supporting shared infrastructure across multiple products or engineering organizations Compensation: •$190,000 - $280,000 / year •Pay within range listed + Bonus + Benefits + Equity Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!