Audio Inference Engineer, Model Efficiency

Cohere - New York, NY

Hiring: Audio Inference Engineer, Model Efficiency Company: Cohere Location: New York, NY Job Posted Time: 2026-09-10 10:17:39 Employment Type: Full-time / Remote Target Skills & Keywords : C++, CV, Deep Learning, LLM, Machine Learning, PyTorch, Python, TensorFlow, Transformers, vLLM About the job Required Skills: •You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Qualifications: •Significant experience developing high-performance audio or machine learning inference systems. •Proficiency with programming languages such as C++ and Python. •Applied hands-on capability in deep learning models for audio, speech, or language applications. •A bias for action and a strong results-oriented mindset. •GPU programming, low-level system optimization, model parallelization techniques over multiple GPUs •Have experience with duplex real-time streaming architectures. •Internals of machine learning frameworks for audio (such as PyTorch, TensorFlow, or specialized audio libraries). •Have experience with inference framework like vLLM, SGLang, Tensort-LLM, or custom distributed inference systems •Sequence modeling (e.g., transformers for audio/speech) and end-to-end audio pipeline optimization •Full-Time Employees At Cohere Enjoy These Perks Compensation: •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!