Principal Software Engineer, Inference

Hewlett Packard Enterprise - Fort Collins, CO

Hiring: Principal Software Engineer, Inference Company: Hewlett Packard Enterprise Location: Fort Collins, CO Job Posted Time: 2026-09-17 05:14:27 Employment Type: Remote Target Skills & Keywords : Accessibility, C++, HTML, Kubernetes, LLM, Nim, Python, RAG, TensorRT, vLLM About the job Experience: •12 years of experience in Software Engineering, including +1 years working directly on LLM inference runtimes or production model serving Required Skills: •Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution •Partner with inference performance engineering teams, with accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency •Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, KV cache offload across GPU memory, host memory, and RDMA-attached storage •Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether each runtime is adopted, developed in-house, or declined •Define the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling •Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences •Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals •Comprehensive understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding Qualifications: •Upstream contribution to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe •Disaggregated prefill/decode serving, or KV cache offload and reuse at scale •RDMA, GPUDirect Storage, InfiniBand, or RoCE •MIG, fractional GPU allocation, and multi-tenant GPU isolation •On-premises, air-gapped, or regulated enterprise software delivery •Minimum of 12 years of experience in Software Engineering, including +1 years working directly on LLM inference runtimes or production model serving •Degree in Computer Science or related field •HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here. •Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability. •What We Can Offer You Compensation: •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!