Staff AI Product Engineer
Nscale - Houston, TX
Hiring: Staff AI Product Engineer Company: Nscale Location: Houston, TX Job Posted Time: 2026-09-16 07:37:33 Target Skills & Keywords : Fine-tuning, Kubernetes, LLM, OpenAPI, PyTorch, Python, Reinforcement Learning, TensorRT, Transformers, Triton, vLLM About the job Experience: •12 years of engineering experience, with significant depth in production AI systems at scale (AI labs, hyperscalers, or leading ML infrastructure companies) Required Skills: •Nscale is looking for a •To set technical direction for the inference and reinforcement learning systems at the core of our AI services platform — and for the APIs through which other engineers consume them. •Set technical direction for Nscale’s inference serving architecture: request routing, scheduling, continuous batching, KV cache management, prefix caching, and speculative decoding •Strategically drive strategy for model efficiency in production — quantization (FP8, INT8/4), sparsity, pruning, distillation, and MoE serving — and the trade-offs between cost, latency, throughput, and model quality •Lead the resolution of systemic performance and reliability challenges across the serving stack, from kernel-level bottlenecks to fleet-level capacity and multi-tenant isolation •Own the architecture of Nscale’s RL and post-training systems: RLHF, DPO/GRPO-style methods, reward modelling, and agentic RL with tool calling, off-policy training, and decoupled sampling and policy updates •Define how inference and training share infrastructure in RL loops — rollout generation, sample buffering, weight synchronization, and the serving engine’s role inside the training system •Establish standards for fine-tuning services (LoRA, QLoRA, adapters, full fine-tuning) and the data curation and processing workflows that feed them Qualifications: •8–12 years of engineering experience, with significant depth in production AI systems at scale (AI labs, hyperscalers, or leading ML infrastructure companies) •Demonstrated ability to set technical direction for an AI systems domain at scale •Deep expertise in production LLM inference: serving architectures, KV cache and memory management, batching and scheduling strategies, speculative decoding, and low-precision inference •Strong hands-on expertise in RL for LLMs — RLHF, DPO and related preference-optimization methods, reward modelling, or agentic/multi-turn RL — including the systems that make them run efficiently on GPU clusters •Proven ability to design developer-facing APIs and SDKs that are clean, versioned, and adopted by other engineers and external customers •Strong cross-team influence: track record of creating standards and practices adopted by multiple teams •Strong proficiency in Python and PyTorch, with a track record of building maintainable, well-tested, production-grade ML systems •Comprehensive expertise in transformer and LLM architectures and their behavior under production load •Demonstrated capacity to design architecture that is both technically excellent and practically adoptable across a diverse engineering organization •Contributions to widely-used open-source inference or RL frameworks (vLLM, SGLang, TensorRT-LLM, verl, OpenRLHF, TRL, DeepSpeed, etc.) Compensation: •Competitive benefits and rewards package •In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!