Machine Learning Engineer (Evals and Voice Models)

Aircall - San Francisco, CA

Hiring: Machine Learning Engineer (Evals and Voice Models) Company: Aircall Location: San Francisco, CA Job Posted Time: 2026-09-16 10:29:34 Employment Type: Contract Target Skills & Keywords : AI, Data Pipeline, Deep Learning, Fine-tuning, LLM, Machine Learning, RAG, Rollup, Systems Design About the job Experience: •3+ years of experience in ML Engineering or Applied ML with 8+ years of overall experience Required Skills: •Design and document comprehensive evaluation frameworks for Aircall’s AI agents across voice, chat and messaging. •Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data, iterating on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency. •Assess AI generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies. •Analyze system design decisions and identify strengths, weaknesses, and potential failure points. •Design annotation guidelines and workflows for human-labeled evaluation data, and calibrate LLM-as-judge systems against human raters to ensure automated evals stay trustworthy over time. •Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals in production, and flagging model or data drift before it impacts customers. •Own the metric contract for every published AI metrics, including definition, population, grain, rollup, validity window. •Build release gates, the offline regression suite each AI surface must pass before a prompt, model, or config change ships, measuring reliability across repeated trials, not just average pass rates. Qualifications: •BS in Computer Science, Machine Learning, Statistics, or related field •Strong experience in evaluating supervised, unsupervised, LLMs and deep learning models. •Applied hands-on capability in failure analysis and evaluating LLMs •Strong communication skills to articulate complex technical concepts across technical and non-technical audiences •Hands-on experience training or fine-tuning voice/speech models (TTS, ASR, or speech-to-speech), including data pipeline construction and experimentation. •MS / PhD in Computer Science, Machine Learning, Statistics, or related field •Base salary range:: $181,000 USD - $250,000 USD Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!