Senior Software Engineer, Inference

Hewlett Packard Enterprise - Durham, NC

Hiring: Senior Software Engineer, Inference Company: Hewlett Packard Enterprise Location: Durham, NC Job Posted Time: 2026-09-17 05:14:27 Employment Type: Remote Target Skills & Keywords : Accessibility, C++, HTML, Kubernetes, LLM, Nim, Python, RAG, TensorRT, vLLM About the job Experience: •8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving Required Skills: •; however, remote work options will be considered. •Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution •Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency •Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage •Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption •Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling •Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence •Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team Qualifications: •Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe •Disaggregated prefill/decode serving, or KV cache offload and reuse at scale •RDMA, GPUDirect Storage, InfiniBand, or RoCE •MIG, fractional GPU allocation, and multi-tenant GPU isolation •On-premises, air-gapped, or regulated enterprise software delivery •Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving •Degree in Computer Science or related field •HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here. •Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability. •What We Can Offer You Compensation: •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!