SRO Lead
Versant Media - Englewood Cliffs, NJ
Hiring: SRO Lead Company: Versant Media Location: Englewood Cliffs, NJ Job Posted Time: 2026-09-03 11:15:43 Target Skills & Keywords : AWS, Azure, CI/CD, CloudFormation, Docker, E2E Testing, GCP, Infrastructure as Code, Integration Testing, Kubernetes, Load Testing, Performance Testing, Root Cause Analysis, Systems Design, Systems Engineering, Terraform About the job Experience: •5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Infrastructure roles. Required Skills: •The System Reliability Engineering (SRE) Lead is a hands-on technical leader responsible for improving the reliability, performance, and scalability of VERSANT’s software, production, and platform systems. •Reporting to the VP of Infrastructure, this role works closely with Software Engineering, Production Engineering, Platform Engineering, and Infrastructure teams to implement reliability best practices, drive end-to-end testing, and ensure systems perform under real-world conditions. •This is a player-coach role focused on execution—building testing frameworks, improving observability, and helping teams proactively identify and resolve system weaknesses before they impact production. •Partner with engineering teams to improve system reliability, availability, and performance. •Help define and implement SLIs, SLOs, and basic reliability standards across services. •Identify reliability gaps and work with teams to address risks in system design and operations. •Contribute directly to code, tooling, and automation that improves system resilience. •Design and implement end-to-end (E2E) testing workflows across distributed systems. Qualifications: •Strong hands-on experience operating and troubleshooting production systems. •Operational familiarity with observability tools (metrics, logging, tracing) and monitoring systems. •Strong debugging and problem-solving skills in complex systems. •Operational familiarity with high-throughput or low-latency systems. •Exposure to SRE concepts such as SLIs/SLOs and incident management practices. •Operational familiarity with infrastructure as code (Terraform, CloudFormation). •Strong collaboration skills and ability to work across multiple engineering teams. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!