Senior / Staff AI Engineer

Snorkel AI - San Francisco, CA

Hiring: Senior / Staff AI Engineer Company: Snorkel AI Location: San Francisco, CA Job Posted Time: 2026-09-10 07:37:32 Target Skills & Keywords : AWS, Airflow, Dagster, Design System, Kubernetes, LLM, Prefect, Python, Reinforcement Learning, Serverless About the job Experience: •5+ years building production software systems, with experience in AI/ML infrastructure, ML platforms, distributed systems, data platforms, or backend infrastructure Required Skills: •You'll work on systems where correctness is not defined by a single deterministic output. Instead, you'll build the infrastructure needed to understand behavior across models, prompts, tools, environments, and multi-step trajectories, and to continuously improve those systems through experimentation and evaluation. •Design and build infrastructure for running large-scale agentic workloads, including multi-step agents interacting with tools, external services, sandboxes, and simulated environments •Build scalable synthetic data generation and automated labeling systems that allow teams to create, refine, and evaluate high-quality training and evaluation datasets •Design evaluation infrastructure for measuring AI system behavior across models, prompts, tools, environments, and multi-step trajectories - including reproducible experiments, benchmark execution, regression detection, and continuous evaluation •Build orchestration and distributed compute systems for running thousands to millions of AI experiments and simulations reliably across heterogeneous compute environments •Develop infrastructure for agent simulation environments, including environment provisioning, isolation, lifecycle management, and scalable execution •Build and operate LLM infrastructure for routing, rate limiting, retries, caching, provider failover, cost attribution, and efficient execution across multiple model providers •Instrument agent and model workloads so failures are observable and debuggable - capturing traces, model interactions, tool calls, environment state, evaluation results, latency, reliability, and cost Qualifications: •Strong proficiency in Python and experience building production-quality APIs, services, and developer tooling •Strong background in distributed systems and cloud platforms (AWS preferred), including compute orchestration, storage, networking, isolation, and failure handling •In-depth knowledge of production system fundamentals - observability, telemetry, reliability, performance, debugging, incident response, and cost management •Demonstrated capacity to reason about AI system quality beyond traditional service metrics, including evaluation design, experiment reproducibility, behavioral regressions, and model or agent variability •Track record of leading complex engineering initiatives, influencing stakeholders, and delivering measurable impact •Demonstrated capacity to work in a fast-paced environment with strong technical communication skills •Fluency with modern AI and developer tooling and a willingness to rapidly evaluate and adopt new models, frameworks, infrastructure, and techniques as the ecosystem evolves Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!