Product Lead, AI/ML (Evals)

Abridge - San Francisco, CA

Hiring: Product Lead, AI/ML (Evals) Company: Abridge Location: San Francisco, CA Job Posted Time: 2026-09-16 12:27:45 Employment Type: Full-time Target Skills & Keywords : Clinical Documentation, Data Pipeline, EMR, LLM, Product Management, Product Strategy About the job Experience: •5+ years of product management experience with significant ownership of ML powered products or platform systems. •5 years of employment. Required Skills: •Abridge's has multiple products - the core product is ambient clinical notes, but we’ve expanded surfaces: billing, clinical decision support, orders, nursing, etc. Measuring quality reliably and iterating fast without breaking clinician trust sit at the core of every single product’s success. •The Evals team owns the tooling, templates, and consultation that product teams use across the eval lifecycle: curate, build, iterate, and deploy. This builds the strategy, platform, and process that enables evals to be fast, trustworthy, consistent, and a true moat. •Drive product strategy and execution for the evals platform. Own the roadmap across the eval lifecycle and own outcomes against it. •Build the shared measurement infrastructure. Help build the systems that let any pod define quality, run experiments, compare models, and watch production. Own the standards for LLM judges, rule-based evaluators, human annotation, and online monitoring, and be clear about where each belongs. •Make model selection fast and routine. Give teams a repeatable way to evaluate a new frontier model within days of release, and a defensible framework for when a post-trained specialty model is worth it over a prompted frontier one. Keep the eval system model-agnostic so it stays a neutral referee. •Own the operating model across pods. Define the eval gates from early build through GA and steady-state monitoring, which gates are hard versus advisory, and who owns non-negotiable floors like critical-error rates. Land this across pods you don't own, without formal authority. •Cross functional execution. Work with Data Engineering on the de-identification pipeline, Data Science on bootstrapping judge quality with less human annotation, Clinical Science on flagged production cases, and the agent platform team as workflows go agentic. •Operate with a high bar for quality, speed, and accountability. Qualifications: •Comprehensive expertise in how to measure and improve model quality, including evaluation frameworks, annotation pipelines, and benchmark design. •Strong technical fluency across ML, data pipelines, and distributed systems. •Demonstrated capacity to balance long term architectural investments with near term quality improvements. •Strong communication skills and the ability to translate complex technical concepts into clear decisions and narratives. •A track record of delivering high quality products in domains where accuracy, reliability, and trust are paramount. •You have experience building evaluation platforms, ML observability systems, or quality measurement pipelines. •You have worked in clinical, healthcare, or regulated environments with a high bar for accuracy and compliance. •You have worked on specialty specific or domain specific model adaptations. •You have worked on personalization systems, context ingestion frameworks, or ambient intelligence products. •You have experience shipping large scale ML products with human in the loop workflows. Compensation: •$250,000 - $290,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!