Senior Machine Learning Operations Engineer

BetMGM - Indiana, United States

Hiring: Senior Machine Learning Operations Engineer Company: BetMGM Location: Indiana, United States Job Posted Time: 2026-09-14 15:10:52 Target Skills & Keywords : API Gateway, AWS, CI/CD, Docker, ECS, Feature Store, Fine-tuning, Flink, GitHub, IAM, IaC, Kafka, Kubernetes, LLM, Lambda, MLOps, Machine Learning, Model Registry, Pinecone, Python, RAG, RBAC, S3, SageMaker, Snowflake, Terraform, VPC, dbt, pgvector About the job Experience: •5+ years shipping software in production — Python, Docker, Kubernetes or ECS, CI/CD, distributed systems debugging — including time on-call. •3+ years operating ML in production — you have owned a model in prod that served real traffic, with stated latency and cost budgets and a runbook you wrote. Required Skills: •Stand up and operate BetMGM's ML platform on AWS (SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (Snowpark ML, Cortex), with Terraform-managed infrastructure. •Build self-service scaffolds that let data scientists ship a model end-to-end without a ticket queue — cookie-cutter project templates with CI, drift monitoring, alerting, IaC, and Snowflake connectivity pre-baked. •Design and operate batch scoring pipelines — SageMaker Batch Transform, dbt-orchestrated scoring against Snowflake, Snowpark ML — with explicit freshness and cost SLAs. •Design and operate real-time inference paths — SageMaker real-time endpoints, Lambda + Bedrock for GenAI, API Gateway — with stated latency budgets (typically sub-100ms) and graceful degradation under load. •Own the feature store (SageMaker Feature Store, Tecton, or Feast) with guaranteed online/offline parity — training-serving skew is treated as an incident, not a tradeoff. •Build CI/CD for ML — model registry, automated retraining triggers, model versioning, lineage from feature → training run → deployed model → live prediction. •Implement champion/challenger, shadow deployments, and canary releases as platform primitives so individual model teams do not reinvent them per project. •Stand up drift detection, data quality, and model performance monitoring (Evidently, Arize, or SageMaker Model Monitor — pick one and standardize) with paging that routes to humans who can fix it. Qualifications: •BS or MS in Computer Science, Math, Statistics, Machine Learning, or other STEM field — or equivalent practical experience. Practical experience wins ties; a PhD is neither required nor a tiebreaker. •AWS depth across the SageMaker surface (Training, Endpoints, Batch Transform, Model Registry, Pipelines) plus the supporting cast (IAM, Lambda, ECS, S3, Secrets Manager, VPC). •Snowflake fluency — Snowpark ML, Cortex, dbt-orchestrated batch scoring, RBAC for ML workloads. •IaC for ML — Terraform + SageMaker Pipelines or equivalent. No manual console deployments to production. •Feature store experience — SageMaker Feature Store, Tecton, or Feast — with explicit ownership of online/offline parity. •Champion/challenger, shadow, and canary deployment patterns as production muscle, not blog-post familiarity. •Drift and model monitoring — Evidently, Arize, WhyLabs, or SageMaker Model Monitor — wired to a paging path. •Software-engineering-first mindset — you treat ML systems as systems, not notebooks. •GenAI in production — Bedrock, Anthropic, or OpenAI APIs integrated into live systems; RAG pipelines; vector DBs (Snowflake Cortex Search, pgvector, Pinecone); evaluation frameworks (Langfuse or in-house). •Snowflake-native ML — Snowpark Container Services, Cortex AISQL, Cortex Agents — for workloads that do not need to leave the warehouse. Compensation: •Medical, Dental, Vision, Life, and Disability Insurance Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!