Staff Site Reliability Engineer
Crunchyroll - San Francisco, CA
Hiring: Staff Site Reliability Engineer Company: Crunchyroll Location: San Francisco, CA Job Posted Time: 2026-09-09 18:01:43 Target Skills & Keywords : Cloud Native, IaC, Infrastructure as Code, Kubernetes, Move, Penetration Testing, TestNG About the job Required Skills: •We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets. •The ideal candidate combines deep expertise in large-scale distributed systems with a strong sense of ownership, collaboration, and service leadership. You are passionate about building highly reliable platforms, eliminating operational toil through automation, and enabling engineering teams to move quickly and safely. In addition, you will champion SecOps best practices by driving vulnerability management, supporting penetration testing initiatives, improving security observability, strengthening cloud and Kubernetes security controls, and ensuring operational readiness for emerging threats. This is a unique opportunity to shape reliability and security engineering practices across CDI while helping build a world-class data and insights ecosystem that enables informed decision-making throughout Crunchyroll. •Core Areas of Responsibility •Reliability Engineering: Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets. •Operational Excellence: Establish and drive best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams. •Observability & Monitoring: Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution. •Automation: Identify operational inefficiencies and develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity. •Platform Scalability: Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency. •Infrastructure Engineering: Drive Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility. •Capacity Planning & Performance: Lead capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently. •Disaster Recovery & Resilience: Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency. •Security Operations (SecOps): Partner with Crunchyroll's security team to integrate security controls, Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!