Principal Site Reliability Engineer (ADEM)
Palo Alto Networks - Santa Clara, CA
Hiring: Principal Site Reliability Engineer (ADEM) Company: Palo Alto Networks Location: Santa Clara, CA Job Posted Time: 2026-09-16 21:54:44 Employment Type: Full-time Target Skills & Keywords : AWS, Ansible, Apache, ArgoCD, Bash, CI/CD, Configuration Management, Docker, EKS, FedRAMP, GCP, GitHub, GitLab CI, GitOps, Grafana, Helm, IaC, Infrastructure as Code, Java, Kafka, Kubernetes, Linux, MySQL, Prometheus, Pulsar, Python, Root Cause Analysis, Systems Engineering, Terraform, Vault About the job Experience: •Must be a U.S Citizen due to Federal Government requirements. •7+ years as an engineer in Infrastructure, Operations, DevOps, or System Engineering. •The candidate must be familiar with and demonstrate proficiency in using code assist and AI productivity tools such as Claude code, Cursor, Windsurf, or GitHub Copilot to accelerate development and troubleshooting. •Expertise in building high-availability, scalable cloud-native applications on GCP (preferred) or AWS. •Expertise in configuration management and IaC (Terraform, Helm, Ansible). Required Skills: •Palo Alto Networks runs a large infrastructure and is one of the largest GCP customers. •Our Infrastructure Platform stack includes Terraform, Kubernetes, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Docker, GCP, AWS, Vault, Kafka, MySQL, Python, Bash, and Go. •Strategically drive success of SRE and DevOps through expert contributions in CI/CD and AIOps initiatives, moving the organization toward self-healing infrastructure. •Architect "Golden Paths" for service delivery, ensuring that SLOs, error budgets, and automated canary analysis are integrated by default. •Design, build, and operate reliable, secure Cloud infrastructure that supports high-scale synthetic monitoring and Real User Monitoring (RUM). •Ensure applications are production-ready, scalable, and resilient, collaborating closely with developers, researchers, and data scientists. •Develop tools and automation frameworks that champion Infrastructure as Code (IaC) and Monitoring as Code (MaC). •Lead root cause analysis (RCA) of critical business and production issues, driving improvements that prevent recurrence. Compensation: •$152,300 - $246,400 / year •Flexible work environment (work from home / hybrid options) •The offered compensation may also include restricted stock units and a bonus Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!