Inference Optimization Engineer

Modular, a Qualcomm company - United States

Hiring: Inference Optimization Engineer Company: Modular, a Qualcomm company Location: United States Job Posted Time: 2026-09-17 01:59:21 Target Skills & Keywords : ASIC, Agile, AutoCAD, CAD, Cloud Native, Kubernetes, LLM About the job Experience: •5+ years of experience in distributed systems or performance engineering. Required Skills: •Candidates based in the US or Canada are welcome to apply. You can work in our office in Los Altos, CA or remotely from home. Onboarding for new hires is conducted in-person in our Los Altos, CA office. •Build the optimization platform that drives inference performance of LLMs served on Modular Cloud to state of the art levels across the latest GPU and ASIC architectures. •Shape the technical direction of Modular Cloud, delivering LLM performance on the Pareto frontier for agentic use cases and keeping it there as the landscape evolves. •Partner closely with the GTM team to deliver highly customized LLM inference tuned to specific customer use cases, and collaborate across engineering to drive optimizations spanning the full stack, from GPU kernels to cloud infrastructure. Translate insights from customer engagements into technical direction for engineering teams. •Publish blog posts on innovative approaches to LLM inference optimization that shape industry wide best practices. •5+ years of experience in distributed systems or performance engineering. •A track record of building durable, reusable software tools and libraries that are adopted across teams and functions. •Sound judgment in evaluating technical tradeoffs and setting priorities, paired with strong communication and technical leadership skills. Compensation: •$198,000 - $286,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!