Site Reliability Engineer (SRE)

Air Apps - San Francisco, CA

Hiring: Site Reliability Engineer (SRE) Company: Air Apps Location: San Francisco, CA Job Posted Time: 2026-09-16 10:59:36 Target Skills & Keywords : AWS, Azure, Bash, CI/CD, CloudFormation, Datadog, Docker, ELK Stack, GCP, Grafana, Helm, IaC, Infrastructure as Code, Kubernetes, Linux, Load Balancing, New Relic, Prometheus, Pulumi, Python, Root Cause Analysis, Systems Design, Systems Engineering, Terraform About the job Experience: •4+ years of experience in Site Reliability Engineering (SRE), DevOps, or System Engineering. Required Skills: •As a Site Reliability Engineer (SRE) •At Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You will work at the intersection of software development and operations, implementing automation, monitoring, and performance optimization strategies to minimize downtime and improve system resilience. •Design and implement scalable, reliable, and fault-tolerant systems across cloud environments. •Develop and maintain observability tools, including monitoring, logging, and alerting (e.g., Prometheus, Grafana, Datadog, ELK). •Automate infrastructure provisioning, deployment, and incident response using Infrastructure as Code (IaC) tools like Terraform or CloudFormation. •Optimize system performance, scalability, and incident response workflows to improve uptime. •Work closely with development and DevOps teams to improve system design for reliability. •Conduct root cause analysis (RCA) and implement preventative measures to minimize failures. Qualifications: •Around 4+ years of experience in Site Reliability Engineering (SRE), DevOps, or System Engineering. •Strong knowledge of cloud platforms (AWS, Azure, or GCP) and cloud-native architectures. •Proficiency in Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Pulumi. •Applied hands-on capability in containerization and orchestration (Docker, Kubernetes, Helm). •Strong Linux system administration and networking fundamentals. •Proficiency in scripting (Bash, Python, or Go) for automation and system monitoring. •Knowledge of load balancing, failover strategies, and distributed systems. •Understanding of security best practices, access control, and compliance requirements. •Strong communication skills and the ability to collaborate with cross-functional teams. •What benefits are we offering? Compensation: •$116,000 - $200,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!