Research Staff, Voice AI Foundations

Deepgram - Ann Arbor, MI

Hiring: Research Staff, Voice AI Foundations Company: Deepgram Location: Ann Arbor, MI Job Posted Time: 2026-09-16 15:21:07 Target Skills & Keywords : Cloudflare, Data Pipeline, Deep Learning, LLM, Transformers, Triton, Twilio About the job Experience: •000 years of audio and transcribed more than 1 trillion words. There is no organization in the world that understands voice better than Deepgram. Required Skills: •Build next-generation neural audio codecs that achieve extreme, low bit-rate compression and high fidelity reconstruction across a world-scale corpus of general audio. •Pioneer steerable generative models that can synthesize the full diversity of human speech from the codec latent representation, from casual conversation to highly emotional expression to complex multi-speaker scenarios with environmental noise and overlapping speech. •Develop embedding systems that cleanly factorize the codec latent space into interpretable dimensions of speaker, content, style, environment, and channel effects -- enabling precise control over each aspect and the ability to massively amplify an existing seed dataset through “latent recombination”. •Design model architectures, training schemes, and inference algorithms that are adapted for hardware at the bare metal enabling cost efficient training on billion-hour datasets and powering real-time inference for hundreds of millions of concurrent conversations. •See "unsolved" problems as opportunities to pioneer entirely new approaches •Can identify the one critical experiment that will validate or kill an idea in days, not months •Have the vision to scale successful proofs-of-concept 100x •Are obsessed with using AI to automate and amplify your own impact Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!