Research Scientist - Generative Audio

Spotify - New York, NY

Hiring: Research Scientist - Generative Audio Company: Spotify Location: New York, NY Job Posted Time: 2026-09-16 16:56:38 Employment Type: Full-time Target Skills & Keywords: Generative Audio, Diffusion Models, Flow Matching, Vocal Synthesis, Speech Synthesis, Post-Training, DPO, RLHF, KTO, PPO, GRPO, Reward Modeling, Reinforcement Learning, Music Generation, Audio Editing, Audio-to-Audio Generation, Text-Guided Music Editing, Machine Learning, Music Information Retrieval, Signal Processing, Probabilistic Modeling, Python, PyTorch, NumPy Experience: - Ph.D. in Computer Science, Mathematics, Engineering, or a related field - Previous industry experience is helpful - Experience in one or more of the following fields: generative modeling, machine learning, music information retrieval, speech processing, audio processing, signal processing, probabilistic modeling, computer vision, or related areas - Deep expertise in at least one focus area: vocal/speech synthesis, post-training alignment techniques (e.g., PPO, GRPO, DPO), or audio-to-audio generation and text-guided music editing - Publications at leading conferences such as ICASSP, ISMIR, INTERSPEECH, ICLR, AAAI, IJCAI, NeurIPS, ICML, CVPR, ECCV, ICCV, or related venues Required Skills: - Conduct groundbreaking research in generative audio using diffusion or flow matching models - Run large-scale experiments using extensive infrastructure - Create practical applications harnessing generative technologies - Collaborate within a cross-functional team of scientists, engineers, product managers, designers, user researchers, and analysts - Publish findings, deliver talks, and attend top conferences - Strong coding skills in Python, PyTorch, and NumPy - Ability to explain complex topics in simple terms and build strong relationships with colleagues and stakeholders Qualifications: - Creative problem solver passionate about building outstanding products that add real value to millions of people - Enthusiastic about turning research ideas into products operating at scale - Focus areas include: Vocal Synthesis (vocal and speech synthesis, ML-based audio processing, signal processing), Post-Training (preference alignment methods such as DPO, RLHF, KTO, reward model design and training, reinforcement learning), Editing (iterative music generation and audio editing, stem replacement, instrumentation change, mood changes, tempo changes, structure changes) - Ability to work within the North Americas region - Team operates within Eastern Standard time zone; core working hours are CET 3pm-6pm / EST 9am-12pm Compensation: - United States base range: $133,194 - $190,278 plus equity - Benefits include health insurance, six month paid parental leave, 401(k) retirement plan, a monthly meal allowance, 23 paid days off, 13 paid flexible holidays Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!