Lead Applied AI Site Reliability Engineer II
Deloitte - Nashville, TN
Hiring: Lead Applied AI Site Reliability Engineer II Company: Deloitte Location: Nashville, TN Job Posted Time: 2026-09-17 06:54:50 Target Skills & Keywords : .NET, AWS, ArgoCD, Azure, Bash, C#, CI/CD, CloudWatch, Datadog, Docker, Dynatrace, Embedded Systems, GCP, GitHub, Grafana, JMeter, Java, Kubernetes, LLM, MLOps, MLflow, Machine Learning, OpenTelemetry, Performance Testing, Prometheus, Python, RBAC, REST, SQL, SonarQube, Splunk, Terraform, Vertex AI, k6 About the job Experience: •6+ years of software engineering and site reliability engineering experience operating large-scale, distributed, cloud-native systems in production, with experience in most of the following: Python, Go, Bash, Java, C#/.NET, SQL/NoSQL, Kubernetes, Terraform, ArgoCD, as well as CI/CD and observability stacks. •3+ years of experience with cloud-native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP-including their AI/ML services such as Azure OpenAI, AWS Bedrock, or Vertex AI-plus container orchestration (Kubernetes, Docker), infrastructure-as-code, networking, and multi-environment management. •1+ years of experience establishing reliability and operational standards-SLO discipline, runbooks, and performance and resilience budgets-including actively leading, mentoring, and guiding team members in the adoption and continuous improvement of these standards. Required Skills: •Excellent interpersonal and organizational skills, with the ability to handle diverse situations, complex projects, and changing priorities, behaving with passion, empathy, and care. Qualifications: •A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline. Experience is the most relevant factor. •Prior experience operating AI/ML and agentic workloads in production-their reliability failure modes (drift, train/serve skew, output variance), MLOps/LLMOps, and the AI control plane (model/LLM gateway, guardrails) from the operability and performance side. •Prior experience with load and performance testing under simulated production traffic (e.g., LoadRunner, k6, or JMeter), chaos engineering (e.g., Azure Chaos Studio, AWS Fault Injector), capacity planning, autoscaling, and cloud/AI cost engineering (FinOps tooling/dashboards, including GPU/inference and token cost attribution). •Prior software engineering experience with the understanding of Business Context Diagrams (BCD), sequence/activity/state/entity relationship/data flow diagrams, OOP/OOD, data structures, algorithms, and code instrumentations, and AI-augmented spec-driven development. •Prior experience using methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks (e.g. LangFuse, LangSmith, or equivalent multi-agent orchestration tools) etc. to operate high-quality, resilient platforms and products at scale. •Candidates must be located within a commutable distance to one of the select locations available for this role •Demonstrated capacity to work in your local office at a minimum of 3 days per week •Demonstrated capacity to travel 10%, on average, based on the work you do and products you build. •Limited immigration sponsorship may be available. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!