Director, Site Reliability Engineering & Service Enablement
ServiceNow - Santa Clara, CA
Hiring: Director, Site Reliability Engineering & Service Enablement Company: ServiceNow Location: Santa Clara, CA Job Posted Time: 2026-09-16 19:09:03 Target Skills & Keywords : ServiceNow About the job Required Skills: •Our Site Reliability Engineering (SRE) team consists of highly skilled engineers responsible for maintaining and enhancing the reliability, scalability, and performance of the ServiceNow infrastructure. Our SRE’s are empowered to resolve technical issues across the entire technology stack, from hardware to applications. Additionally, they work to improve the platform's operability, aiming to reduce the number of incidents and minimize Mean Time to Recovery (MTTR). To achieve this, the team combines software development, networking, database, and systems engineering skills to tackle complex problems, striving to maintain our platform operating for our customers. •We are looking for a Director of Site Reliability Engineering to lead the next phase of our reliability transformation as ServiceNow modernizes toward a cloud-agnostic, cloud-ready production platform. •This leader will own key elements of the SRE operating model across •Reliability Engineering, Service Enablement, Service Registry, SLI/SLO standards, reliability governance, automation, AI-enabled operations, and production readiness •. The role will lead a global engineering organization and partner across Product Engineering, Infrastructure, Architecture, Security, Release Engineering, and Customer Support to establish consistent reliability practices across ServiceNow products and services. •The Director will play a critical role in evolving the organization from reactive operations toward an •engineering-led SRE model focused on prevention, automation, resilience, and continuous improvement •What You Get To Do In This Role •Define and execute the SRE strategy and operating model across reliability engineering, service enablement, observability, automation, incident learning, and production readiness. •Lead and develop a global organization of engineering managers, technical leaders, and SREs. •Establish enterprise reliability standards for service ownership, tiering, golden signals, SLIs/SLOs, error budgets, alerting, on-call practices, and service health reviews. •Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services. •Own the Service Registry strategy, improving service ownership, dependency visibility, maturity tracking, and impact-aware operational decision-making. •Drive adoption of SLIs, SLOs, error budgets, and burn-rate alerting across critical services, ensuring teams consistently use reliability signals to manage customer impact. •Build a culture of engineering away toil by turning recurring operational work and incident patterns into automation, self-service, and systemic fixes. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!