Staff Site Reliability Engineer (SRE)
EarnIn - Mountain View, CA
Hiring: Staff Site Reliability Engineer (SRE) Company: EarnIn Location: Mountain View, CA Job Posted Time: 2026-09-17 00:54:44 Employment Type: Full-time / Hybrid Target Skills & Keywords : AWS, CloudWatch, Datadog, DynamoDB, EKS, Kafka, Kubernetes, LLM, OpenTelemetry, Python, RDS, SOC 2, SQS, Slack, Terraform About the job Experience: •7+ years in SRE, Software Engineering, or Infrastructure Engineering with increasing scope and cross-org influence. Track record of KPI driven reliability and operational excellence improvements at scale. Required Skills: •Set a reliability strategy with AI at the center. Define SLIs, SLOs, and error budgets across critical services. Use AI to surface trends, predict capacity risks, and auto-generate reliability scorecards so teams act on data. •Redesign the incident lifecycle around AI-assisted speed. Lead high-severity incident response as IC. Build AI-driven alert correlation and triage that reduces noise and accelerates root-cause identification. Drive adoption of AI-generated postmortems that surface systemic patterns and automatically track corrective actions through to completion. •Improve on-call fundamentally better through automation. Build AI agents that draft runbook responses, pull relevant context from Datadog, incident.io, and Slack during pages, and recommend remediation steps, so on-call engineers spend less time deciding and searching. •Push AI-first operations into product engineering teams. Partner with product engineering to embed AI-assisted investigation, alerting, and production readiness into their workflows. Make AI tooling the default path for every team that owns a service, not an SRE-only capability. •Architect for resilience at scale •Guide service designs for graceful degradation, failure isolation, and capacity planning across EarnIn's AWS footprint (EKS, Kafka, DynamoDB, RDS, SQS). Use AI-driven analysis to identify architectural weak points before they become incidents. •Raise the bar through mentorship and standards. Coach engineers on reliability practices, run design and incident reviews, and build documentation and tooling that makes reliability knowledge accessible. Set the expectation that AI-assisted workflows are how EarnIn operates, not an experiment. •Demonstrated experience improving reliability and operational excellence at scale using clear KPIs such as MTTR, MTTD, alert quality, incident recurrence, SLO attainment, on-call health, or corrective-action completion. Compensation: •$252,000 - $308,000 / year •Flexible work environment (work from home / hybrid options) •The base salary range for this full-time position is $252,000-$308,000, plus equity and benefits Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!