Observability & SRE Engineer
World Wide Technology - New Home, MO
Hiring: Observability & SRE Engineer Company: World Wide Technology Location: New Home, MO Job Posted Time: 2026-09-10 10:38:08 Employment Type: Full-time Target Skills & Keywords : AWS, Accessibility, Agile, Ansible, Azure, CI/CD, Docker, ELK Stack, GCP, Git, GitHub, Grafana, Infrastructure as Code, Kubernetes, Linux, Load Balancing, Machine Learning, OpenTelemetry, Podman, Prometheus, Python, Root Cause Analysis, Splunk, Terraform About the job Experience: •Test-First development mindset with functional, end-to-end and regression testing experience •Hands-on SRE experience with SLOs, error budgets, blameless reviews, toil reduction, and production readiness •Public cloud platform experience (AWS, Azure, or GCP) •Operational familiarity with distributed tracing and log aggregation tooling (e.g., OpenTelemetry, ELK/Splunk) Required Skills: •Collection and strategic application of metrics to drive organizational decisions •Providing a holistic view of system health using observability practices •APM, RUM, and Synthetic Transaction monitoring •Driving reliability through monitoring, alerting, observability, and SRE practices •Applying AI-first thinking to automate workflows, correlate alerts, surface insights, and speed incident response •Improving reliability through SLOs, SLIs, error budgets, incident reviews, and continuous improvement •Defining and maintaining on-call rotations, escalation paths, and incident response processes, including participating in on-call coverage •Reducing operational toil through automation, self-healing systems, and infrastructure as code Qualifications: •5+ years of professional experience in IT Operations •Knowledge of programming languages Python, Go, or equivalent •Knowledge of system frameworks including Git and GitHub •Critical thinker with excellent written and verbal communication skills •Team-oriented individual with very strong work ethic •Operational familiarity with Linux, preferably administrative knowledge •Understanding of container technologies •Understanding of SRE concepts such as SLIs, SLOs, incident response, root cause analysis, capacity planning, and reliability automation •Practical understanding of AI-enabled tools, automation patterns, and responsible AI use to improve operational efficiency •Willingness to participate in an on-call rotation and drive incidents through triage, mitigation, and resolution Compensation: •$90,000 - $135,000 / year •Certain positions may also be eligible for variable incentive compensation, such as bonuses or commissions, that are not included in the base pay Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!