Inference Research Engineer (MTS)

Coral Bricks AI - San Francisco, CA

Hiring: Inference Research Engineer (MTS) Company: Coral Bricks AI Location: San Francisco, CA Job Posted Time: 2026-09-16 14:25:09 Target Skills & Keywords : C++, CV, GitHub, LLM, PyTorch, TensorRT, Triton, vLLM About the job Required Skills: •You'll research and build new ways to make LLM inference faster and cheaper, then prove them against real agent workloads. The work sits between research and systems engineering: form a hypothesis about where time or memory is going, design the experiment, implement the change, and measure whether it survives contact with production-shaped traffic. •This role is distinct from our GPU infrastructure role. You won't own the day-to-day operation of the fleet or deployment platform. You'll own the performance ideas that change what the serving system can do: new scheduling policies, cache strategies, parallelism approaches, quantization methods, and model-specific optimizations. •This is a founding-team role open to all experience levels, including new grads. You'll work close to production, see your changes show up directly in customer cost and latency, and learn the deep end of the stack on the job. •Research ways to push throughput and bring down time-to-first-token and inter-token latency on workloads that are heavy on prompts, long-running, and high-fanout — batching, scheduling, parallelism, caching, and quantization. •Cross-platform serving: we're building for both NVIDIA and AMD GPUs — bring models up on each, close the performance gap between them, and keep the stack portable rather than vendor-locked. •Prototype changes across the GPU and CPU sides of the serving stack when the research requires it, and carry successful ideas far enough to demonstrate them under production-shaped load. •Bring up and tune new open-weight model families as they release. •Build the profiling and load-testing harness that tells us — honestly — what the system is doing under real agent load, not synthetic benchmarks. Qualifications: •Open-source contributions to vLLM, SGLang, TensorRT-LLM, llama.cpp, or similar. Compensation: •$120,000 - $200,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!