Cloud Reliability Engineer

Versant Media - Englewood Cliffs, NJ

Hiring: Cloud Reliability Engineer Company: Versant Media Location: Englewood Cliffs, NJ Job Posted Time: 2026-09-10 11:15:43 Employment Type: Hybrid Target Skills & Keywords : AWS, Bash, CI/CD, CloudFormation, Infrastructure as Code, PowerShell, Python, Root Cause Analysis, Terraform About the job Experience: •7 years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, Infrastructure Engineering, or related roles. Required Skills: •Responsible for ensuring the availability, performance, scalability, and operational excellence of VERSANT’s cloud platforms and services. •As a leading media company, VERSANT operates digital products, streaming platforms, content delivery systems, and media workflows that demand high levels of uptime and performance. The Cloud Reliability Engineer will help ensure these services remain resilient, scalable, and operationally mature. •The ideal candidate has strong experience with AWS, monitoring and observability platforms, incident management, automation, infrastructure as code, and operational best practices. Experience with AWS Organizations, Control Tower, Identity Center, Terraform, and modern cloud operations tooling is highly desirable. •Design, implement, and maintain reliability practices for cloud infrastructure and platform services. •Define and monitor service-level objectives (SLOs), service-level indicators (SLIs), and operational metrics. •Identify reliability risks and implement solutions that improve availability, scalability, and resilience. •Drive continuous improvement initiatives focused on operational excellence and system stability. •Design and maintain monitoring, logging, alerting, and observability solutions across AWS environments. Qualifications: •Bachelor’s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience. •3–7 years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, Infrastructure Engineering, or related roles. •Strong hands-on experience with AWS cloud services and enterprise-scale AWS environments. •Incident management and root cause analysis •Operational troubleshooting and performance tuning •AWS Organizations, Control Tower, and Identity Center •In-depth knowledge of AWS networking, resiliency, and cloud architecture concepts. •Strong troubleshooting, communication, and collaboration skills. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!