Senior AI/ML Engineer
Bain & Company - Atlanta, GA
Hiring: Senior AI/ML Engineer Company: Bain & Company Location: Atlanta, GA Job Posted Time: 2026-09-10 06:32:10 Employment Type: Full-time / Hybrid Target Skills & Keywords : AWS, Assembly, Databricks, Docker, Feature Flags, Feature Store, IAM, Kubernetes, LLM, LangChain, LlamaIndex, MLflow, Machine Learning, Model Registry, OpenTelemetry, Prometheus, Python, RAG, SageMaker, Terraform, TestNG, pgvector, pytest About the job Experience: •6+ years of experience building and operating production ML systems, including model deployment, serving, and post-deployment monitoring. •3 years of service and is 100% vested upon start date Required Skills: •Core ML Systems Deployment, Serving, and Operations (80%) •Build, deploy, and operate production inference and serving systems for models, embeddings, and re-rankers: request batching, concurrency, and throughput tuning against latency and cost SLAs. •Own the model and prompt lifecycle in MLflow: packaging, model registry governance, promotion workflows, staged rollout behind feature flags, and clean rollback. •Build and maintain LLMOps tooling: prompt and instruction versioning, model-gateway configuration (e.g., Portkey), inference orchestration, and response caching and cost controls. •Build and operate production RAG and retrieval pipelines end-to-end: structure-aware chunking, contextual embedding, hybrid vector plus keyword retrieval, cross-encoder re-ranking, and context assembly. •Design and maintain model and retrieval evaluation frameworks: golden datasets, metric definitions, LLM-as-judge with calibration, regression gates in CI, and production drift monitoring. •Instrument production ML systems with structured logs, OpenTelemetry spans, and Prometheus metrics: token usage, latency percentiles, retrieval hit rates, drift, and hallucination monitoring, with dashboards and alerting. •Partner cross-functionally with the Agent / AI squad to serve model and retrieval outputs as structured tool responses consumed by the Agent Gateway; partner with Data Engineers on feature and embedding pipelines. Qualifications: •Bachelor’s degree in Computer Science, Engineering, Machine Learning, Data Science, Statistics, or a related field (or equivalent practical experience). •Demonstrated experience owning the model and prompt lifecycle end-to-end (packaging, registry, promotion, rollout, monitoring, and rollback). •Demonstrated experience building and operating production RAG or retrieval systems end-to-end, from embedding and retrieval through re-ranking and evaluation. •Strong Python for production ML; code written to production standards (testing, linting, typing). •Demonstrated ability to mentor other engineers and raise engineering standards through code review and repository conventions. •Strong Python: ML and serving code written to production engineering standards (type hints, Pydantic, pytest, Ruff, mypy strict). •MLflow: experiment tracking, model registry, custom model flavours, promotion workflows, and model serving configuration. •LLMOps tooling: prompt and instruction versioning, model gateways (e.g., Portkey), inference orchestration frameworks (LangChain, LlamaIndex, or equivalent), and response caching. •Model serving and inference optimisation: request batching, concurrency and throughput tuning, latency budgeting, and awareness of quantisation and hardware trade-offs. •RAG pipeline engineering: chunking strategies, contextual embedding, hybrid retrieval, cross-encoder re-ranking, and context assembly. Compensation: •$128,000 - $153,750 / year •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!