SRE - Senior AI Platform Reliability Engineer
HTC Global Services - Seattle, WA
Hiring: SRE - Senior AI Platform Reliability Engineer Company: HTC Global Services Location: Seattle, WA Job Posted Time: 2026-09-10 12:34:10 Employment Type: Hybrid Target Skills & Keywords : AWS, AppDynamics, Azure, Bash, CI/CD, GCP, Grafana, Helm, Infrastructure as Code, Kafka, Kubernetes, LLM, Machine Learning, MongoDB, OpenTelemetry, PostgreSQL, Prometheus, Python, Redis, Splunk, Terraform, Vault About the job Experience: •7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, cloud infrastructure, or a related field. Required Skills: •The selected engineer will help build, scale, and operate the cloud and Kubernetes infrastructure that enables enterprise AI capabilities across the company and the selected engineer will help design, scale, and operate the cloud and Kubernetes infrastructure supporting enterprise AI workloads. •You Do Not Need To Be An AI Model Developers, Data Scientists, Or LLM Experts. Instead Your Day To Day Focus And Experience With The Infrastructure And Operational Challenges Involved In Deploying AI models or AI services into production •Operating AI platforms at enterprise scale •Supporting AI workloads across cloud environments •Building reliable infrastructure for model inference, APIs, and AI-enabled applications •Monitoring the performance, availability, and capacity of production AI platforms •The ideal candidate combines strong AI platform infrastructure experience with deep expertise in SRE, Kubernetes, multi-cloud engineering, Infrastructure as Code, observability, and production reliability. •Lead the design, implementation, and operation of highly available infrastructure supporting enterprise AI platforms and services. Qualifications: •Direct experience supporting AI, machine learning, or model-serving platforms in production. •In-depth knowledge of the infrastructure required to move AI services from development into production. •Expert-level Kubernetes administration and production operations experience. •Strong Helm and Terraform experience. •Applied hands-on capability in Google Cloud Platform, with additional AWS or Azure experience. •Strong scripting and automation skills using Python, Bash, and YAML. •Strong production troubleshooting and incident-management experience across cloud-native distributed systems. Compensation: •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!