Research Engineer, Audio and Speech
Decagon - San Francisco, CA
Hiring: Research Engineer, Audio and Speech Company: Decagon Location: San Francisco, CA Job Posted Time: 2026-09-10 10:27:00 Employment Type: Full-time Target Skills & Keywords : Machine Learning, PyTorch, Python About the job Experience: •2+ years of experience in speech, audio ML, multimodal ML, or production machine learning Required Skills: •As a Research Engineer focused on Audio and Speech, you’ll be responsible for building the models and agent harnesses that power Decagon’s real-time voice agents and taking them all the way from idea to production. Your work will advance multimodal and full-duplex systems that can listen, reason, speak, and respond naturally in real time. •We’re looking for strong engineers who want to build the next generation of AI voice agents. People here own their work end-to-end, ship real improvements, and are trusted to make high-impact technical decisions. •In this role, you will •Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction •Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech •Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages •Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes •Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale Qualifications: •Operational familiarity with speech-to-speech or full-duplex models Compensation: •$200,000 - $400,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!