Lead Site Reliability Engineer
Lumen Technologies - United States
Hiring: Lead Site Reliability Engineer Company: Lumen Technologies Location: United States Job Posted Time: 2026-09-12 13:33:55 Employment Type: Remote Target Skills & Keywords : AWS, Ansible, Azure, CI/CD, CloudWatch, Datadog, GCP, Grafana, Infrastructure as Code, Kubernetes, Prometheus, Python, Root Cause Analysis, Systems Design, Systems Engineering, Terraform About the job Experience: •8+ years in software development, systems engineering, and/or networking •5+ years of related experience required. Required Skills: •Lumen's Network as a Service (NaaS) platform delivers on-demand networking at scale. As Lead SRE, you'll own the reliability of that platform — partnering with operations teams and development counterparts to drive technical direction and resolve systemic issues across a broad range of network topologies and applications. •Success in this role draws on networking fundamentals, cloud platforms, software development and troubleshooting methodology, and a bias toward automating what you'd otherwise do twice. We're looking for a change maker — someone who sees where the platform should go next and drives meaningful impact for the customers who rely on it •This role is designated as a fully remote position within the United States. •Serve as subject matter expert for network automation platform applications, services, and hosting environments •Build and maintain the observability stack: instrument services, collect and curate metrics, and create dashboards and visualizations that make system health obvious at a glance •Define and tune proactive alerting so issues surface before customers feel them •Champion core SRE principles — SLIs, SLOs, and error budgets — and advocate for resilient, fault tolerant architecture •Incident Management Qualifications: •What We Look For in a Candidate •Bachelor’s degree or equivalent in engineering, computer science, or related field. •Applied hands-on capability in at least one major cloud platform (AWS, Azure, or GCP), including compute, networking, and identity services •Strong automation and infrastructure-as-code skills: Terraform, Ansible, and Python •Solid functional working knowledge of modern observability and monitoring tooling (e.g., Datadog, CloudWatch, Grafana, Prometheus), including building dashboards and defining alerts •Demonstrated experience with incident management and blameless postmortems •Comfort using AI-assisted development and agentic tools as part of daily engineering practice •Understanding of network technologies including Internet, Ethernet, IPVPN, Edge Compute, and Optical transport •Strong listening and communication skills; able to operate with autonomy while knowing when to escalate •Multi-cloud experience across AWS, Azure, and GCP Compensation: •$105,786 - $141,047 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!