Site Reliability Engineering (SRE) Manager
McAfee - Frisco, TX
Hiring: Site Reliability Engineering (SRE) Manager Company: McAfee Location: Frisco, TX Job Posted Time: 2026-09-10 12:18:12 Employment Type: On-site Target Skills & Keywords : AWS, CloudWatch, EKS, GCP, Grafana, Kubernetes, Python, SQL, Stakeholder Management, Terraform About the job Experience: •9+ years of experience in Site Reliability Engineering, DevOps, Infrastructure, or related roles, including significant experience in a leadership or management capacity. Required Skills: •Lead and grow a team of SREs, setting technical direction and reliability strategy across AWS, GCP infrastructure and EKS platforms, and GKE platforms. •Own the organization's Incident and Problem Management processes, ensuring major incidents are handled efficiently, with timely executive communication and thorough post-incident reviews. •Strategically drive team's automation strategy, championing Python-based tooling and frameworks that reduce manual toil and improve reliability at scale. •Set standards for Terraform-based infrastructure-as-code, ensuring secure, scalable, and consistent provisioning practices across teams. •Define the organization's observability strategy, ensuring Grafana dashboards, alerting, and SQL/CloudWatch-based analysis practices scale effectively. •Act as a senior escalation point and incident commander for the most critical, high-severity incidents, providing calm, decisive leadership under pressure. •Partner with senior leadership, product, and engineering stakeholders to communicate risk, reliability posture, and remediation roadmaps clearly and confidently. •Own hiring, mentoring, performance management, and career development for the SRE team. Qualifications: •Demonstrated experience building, leading, and growing high-performing technical teams. •Strategic reliability leadership: Balances operational excellence with long-term reliability improvements, ensuring the team addresses immediate risks while building scalable, sustainable practices. •Deep, hands-on background with AWS infrastructure and strong technical credibility to guide architecture and operational decisions. •Proven track record leading teams through complex EKS/GKE troubleshooting and operational challenges. •Strong technical fluency in Python for automation and Terraform for infrastructure-as-code, with the ability to guide and review the team's work. •Extensive experience owning ITSM processes — Incident and Problem Management — at an organizational level. •Strong command of observability practices, including Grafana dashboards, SQL, CloudWatch Logs Insights, and alerting strategy. •Outstanding communication skills — able to clearly articulate technical risk, incident impact, and strategy to executive leadership and cross-functional stakeholders. •Strong stakeholder management skills, comfortable operating at the intersection of engineering, product, and business leadership. •AWS Certification — required (e.g., AWS Certified Solutions Architect – Professional, AWS Certified DevOps Engineer – Professional, or equivalent). Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!