AI Infrastructure Engineer - Inference Platform
Hoonify Technologies Inc. - Albuquerque, NM
Hiring: AI Infrastructure Engineer - Inference Platform Company: Hoonify Technologies Inc. Location: Albuquerque, NM Job Posted Time: 2026-09-16 16:38:22 Target Skills & Keywords : ASIC, Bash, C++, CI/CD, FPGA, Fine-tuning, Git, Grafana, Kubernetes, LLM, Linux, Node.js, Prometheus, Python, RAG, Rust, TensorRT, vLLM About the job Required Skills: •Deploy, tune, and optimize high-performance LLM inference pipelines on GPU infrastructure, improving throughput, latency, and cost efficiency within established design patterns. •Analyze, profile, and optimize model serving workloads across inference frameworks such as vLLM, SGLang, and TensorRT-LLM, and across different model families and hardware architectures. •Build and operate scalable, production-grade API services for model inference, including request routing, multi-tenant isolation, usage metering, and observability. •Develop benchmarking harnesses, monitoring infrastructure, and automation tooling that make serving performance measurable and reproducible. •Scale inference workloads across multi-GPU, multi-node environments spanning NVIDIA and AMD accelerators. •Evaluate, prototype, and integrate model fine-tuning workflows and frameworks. •Collaborate closely with engineering and product teams to align infrastructure capabilities with customer-facing services. •Investigate and resolve issues across the stack, including container, node, network, and accelerator-level problems, escalating appropriately when scope exceeds the role. Qualifications: •Bachelor's degree in Computer Science, Computer Engineering, Applied Math, or Data •Science, plus three (3) years relevant work experience or equivalent combination of education and relevant experience. •Professional software engineering experience, with at least some of it touching ML systems, GPU workloads, or high-performance backend services. •Solid functional working knowledge of Kubernetes in a production context, including writing and debugging manifests, understanding core resource types, and operating production workloads. •Hands-on experience serving or deploying LLMs — you've run vLLM, SGLang, TGI, TensorRT-LLM, or similar. •Comfort working in a Linux environment and with standard developer tooling, including Git-based workflows. •Operational familiarity with CI/CD systems and the basic mechanics of automated build, test, and deployment pipelines. •Strong proficiency in Python with familiarity at least one programming or scripting language used for infrastructure work (Go, Rust, C++ or Bash). •Fine-tuning experience of any depth: LoRA/QLoRA, full fine-tunes, dataset curation, or evaluation design. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!