Sr Site Reliability Engineer
Realtor.com - Austin, TX
Hiring: Sr Site Reliability Engineer Company: Realtor.com Location: Austin, TX Job Posted Time: 2026-09-03 10:41:46 Target Skills & Keywords : API Gateway, AWS, ArgoCD, Bash, CI/CD, CircleCI, CloudFormation, CloudFront, CloudWatch, Datadog, Docker, EC2, ECS, EKS, Fargate, Financial Planning, GDPR, GitHub Actions, GitOps, Go, Grafana, GraphQL, Helm, IAM, IaC, Infrastructure as Code, Istio, Java, Jenkins, Kubernetes, Lambda, OpsGenie, PagerDuty, Prometheus, Python, RDS, S3, SOC 2, Service Mesh, ServiceNow, Splunk, Terraform, User Experience, VPC, Vault About the job Experience: •5+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering with demonstrated success improving system reliability •3+ years hands-on experience with AWS (EKS, EC2, RDS, S3, CloudWatch, IAM) and Kubernetes including cluster management Required Skills: •Implement and maintain highly available AWS infrastructure including EKS clusters, Fargate (ECS), and multi-region architectures •Support reliability of critical services: Skyway (CI/CD), Frontdoor (Tyk), Pantheon (Apollo GraphQL), and supporting infrastructure •Monitor SLIs, SLOs, and error budgets for Tier 1/2/3 systems; participate in architectural reviews for reliability and cost-efficiency •Implement reliability patterns including circuit breakers, graceful degradation, and automated failover •Implement observability solutions using NewRelic for APM, distributed tracing, metrics, and logging for rapid troubleshooting •Build dashboards and alerts that reduce MTTD and MTTR; contribute to observability standards across teams •Identify infrastructure cost optimization opportunities and implement FinOps practices including rightsizing and resource lifecycle management •Support cost-conscious architecture decisions and CI/CD spend optimization (CircleCI, Argo CD) Qualifications: •Bachelor’s degree or equivalent experience •Proficient programming skills (Python, Go, or Java) with infrastructure automation and Infrastructure as Code experience (Terraform, CloudFormation) •Production experience with observability tools (NewRelic, Datadog, Prometheus, Grafana, Splunk) and distributed systems •Preferred: Exposure to chaos engineering tools, API Gateway technologies (Tyk/Kong), GraphQL federation (Apollo), cost optimization initiatives, FinOps principles Compensation: •Immediate eligibility into Company 401(k) plan with 3 Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!