Manager, Site Reliability Engineer

Forge - New York, NY

Hiring: Manager, Site Reliability Engineer Company: Forge Location: New York, NY Job Posted Time: 2026-09-03 14:17:47 Target Skills & Keywords : AI, AWS, Ansible, Azure, CI/CD, CloudWatch, Datadog, Kubernetes, Terraform About the job Experience: •5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function. •10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience. Required Skills: •Manage Forge's Site Reliability Engineering team responsible for keeping Forge systems highly available for customers. •Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post-incident learning. •Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics. •Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation. •Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards. •Contribute to technical design, architecture, automation, infrastructure, and overall team delivery. •Partner cross-functionally with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability. •Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health. Qualifications: •Bachelor's degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience. •Applied hands-on capability in observability, monitoring, alerting, incident response, troubleshooting, and production support. •Strong technical judgment, communication skills, and ability to influence across engineering and non-engineering stakeholders. •Operational familiarity with Kubernetes, container platforms, infrastructure-as-code, Terraform, Ansible, or similar automation tooling. •For residents of San Francisco/Bay Area, CA or New York, NY the annual salary range for this role is $150,000-$220,000 + annual bonus. Final offers may vary from the amount listed based on geography, candidate experience and expertise, annual bonus, and other factors. •Upon offer, we conduct background checks that include employment and education verification, state, and county criminal history searches as well as fingerprint and drug test. Compensation: •$150,000 - $220,000 / year •For residents of San Francisco/Bay Area, CA or New York, NY the annual salary range for this role is $150,000-$220,000 + annual bonus Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!