Staff Site Reliability Engineer - Kubernetes
Okta - Washington, DC
Hiring: Staff Site Reliability Engineer - Kubernetes Company: Okta Location: Washington, DC Job Posted Time: 2026-09-02 23:13:17 Target Skills & Keywords : AWS, Ansible, Bash, CI/CD, CircleCI, CloudFormation, CloudWatch, Compliance, Docker, EC2, ECS, EKS, ELK Stack, Encryption, GitLab, Grafana, Helm, IAM, Istio, Jenkins, Kubernetes, Okta, Prometheus, Python, RBAC, RDS, S3, Service Mesh, Spinnaker, Terraform, TestNG About the job Experience: •Supporting Your Well-Being •Developing Talent and Fostering Connection + Community •We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one. Required Skills: •Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency. •AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement best practices for cost management, scaling, and security within AWS. •Helm Management: Utilize Helm to automate and streamline the deployment of applications and services to Kubernetes clusters. Create, maintain, and manage Helm charts for production-ready deployments. •Karpenter Implementation: Implement and manage Karpenter to dynamically scale Kubernetes clusters in response to workload demands. •Istio Service Mesh Management: Configure and manage Istio to provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained traffic management, service discovery, and policy enforcement. •Platform Automation & Scaling: Automate the deployment, scaling, and management of infrastructure and applications. Work with CI/CD pipelines to ensure a seamless flow from development to production with minimal downtime. •Incident Management & Troubleshooting: Respond to incidents, troubleshoot, and resolve system issues related to performance, availability, and security in a timely and effective manner. •Security & Compliance: Design and implement secure cloud infrastructure with appropriate access controls, network security, and compliance frameworks. Qualifications: •4+ years of experience with Kubernetes/Helm •4+ years of Experience with Terraform. •5+ years of Experience with AWS •Practical experience utilizing multi-region cloud environments. •Proven experience with AWS (EC2, RDS, S3, CloudFormation, IAM, etc.) and solid understanding of cloud-native architectures. •Strong expertise in Kubernetes platform creation, management, and optimisation (e.g., setting up highly available clusters, networking, and storage). •Applied hands-on capability in Helm for Kubernetes application deployment and management. •Practical experience with Karpenter for dynamic scaling of Kubernetes clusters and optimising resource usage. •Expertise in managing and securing Istio for service mesh, including traffic management, security, and observability features. •Proficiency in CI/CD pipelines and automation tools (e.g., Jenkins, GitLab, CircleCI, Terraform, Ansible, Spinnaker). Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!