AI, HPC & GPU Infrastructure Support Engineer
Vast.ai - Los Angeles, CA
Hiring: AI, HPC & GPU Infrastructure Support Engineer Company: Vast.ai Location: Los Angeles, CA Job Posted Time: 2026-09-12 10:29:34 Employment Type: Full-time / On-site Target Skills & Keywords : Bash, CentOS, DNS, Debian, Docker, Docker Compose, Firmware, Grafana, LLM, Linux, Prometheus, PyTorch, Python, RHEL, TensorFlow, Ubuntu, VPN About the job Required Skills: •This role focuses on troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, GPU workloads, Ubuntu, Docker, KVM based virtual machines, networking, hardware, BIOS, and firmware. You’ll investigate failures, reproduce issues, identify root causes, and propose practical solutions across the full infrastructure stack. •You’ll also serve as the engineering resource our L1 support team relies on when tickets go beyond frontline triage. You’ll own complex escalations end-to-end, gather technical evidence, coordinate with the appropriate teams, and communicate findings clearly to clients, infrastructure suppliers, and internal teams. •The best engineers in this role don’t just resolve individual issues—they recognize recurring patterns, improve diagnostic tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic Linux, GPU, and infrastructure problems. •Strong GPU troubleshooting experience, Linux systems knowledge, and technical support skills are the primary requirements. You should be comfortable working autonomously in Ubuntu environments and troubleshooting NVIDIA drivers, CUDA, containers, virtual machines, networking, hardware, and GPU workloads. •Vast.ai users or hosts strongly preferred. •This is a full-time position based in our Westwood, Los Angeles office. •Sunday–Thursday: Four days on-site and one day working from home •Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments Qualifications: •15 minutes — Initial Screening (Virtual): A brief conversation about your background, availability, and interest in the role •45 minutes — Experience Interview (Virtual): An introduction to Vast.ai and a deeper discussion of your technical and support experience •2 hours — Meet and Greet and Technical Assessment (On-site): Meet the team and complete an LLM-assisted Linux systems operations assessment Compensation: •$90,000 - $150,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!