Staff Research Engineer, Model Efficiency
Cohere - New York, NY
Hiring: Staff Research Engineer, Model Efficiency Company: Cohere Location: New York, NY Job Posted Time: 2026-09-10 10:17:39 Employment Type: Full-time / Remote Target Skills & Keywords : CV, LLM, Machine Learning About the job Required Skills: •Model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality •As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. •An appetite to work in a fast-paced high-ambiguity start-up environment •Publications at top-tier conferences and venues (ICLR, ACL, NeurIPS) •Full-Time Employees At Cohere Enjoy These Perks •A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch. •Full health and dental benefits, including a separate budget for mental health. •RRSP matching, 401K, Pension Scheme. Qualifications: •Have a PhD in Machine Learning or a related field •Understand LLM architecture, and how to optimize LLM inference given resource constraints •Have significant experience with one or more techniques that enhance model efficiency Compensation: •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!