Staff AI Infrastructure Engineer
Biohub - Redwood City, CA
Hiring: Staff AI Infrastructure Engineer Company: Biohub Location: Redwood City, CA Job Posted Time: 2026-09-03 10:23:45 Employment Type: Hybrid Target Skills & Keywords : ArgoCD, Bash, C, C++, GitOps, Grafana, Helm, Kubernetes, Linux, Node.js, Prometheus, PyTorch, Python, Rust, Systems Design, TCP/IP About the job Experience: •8+ years of AI/ML infrastructure engineering experience, with deep expertise in at least one of: HPC/Slurm cluster operations, Kubernetes at scale, distributed systems debugging, or GPU compute infrastructure. Required Skills: •Own reliability, observability, and incident response for multi-site GPU clusters running Slurm on Kubernetes. Build the systems, automation, and processes that keep clusters healthy, and that enable fast, efficient recovery when things break. •Debug and resolve deep infrastructure failures across storage, networking, scheduling, and GPU compute layers. Build the tooling and operational patterns that make these failures easier to detect, diagnose, and prevent. •Design and execute GPU cluster scaling plans, systematically validating storage, networking, interconnect, and scheduler behavior as clusters grow to support larger training runs. •Build automation and tooling to manage cluster operations at scale: capacity planning, GPU utilization monitoring workload manager policy management, and pod lifecycle automation. •Drive configuration-as-code practices, ensuring cluster state is reproducible and auditable, and managed through version-controlled pipelines. •Collaborate directly with AI researchers and hero run leads to understand training workload patterns and design infrastructure that meets frontier-scale requirements. •Own the vendor relationship on technical issues — escalating SEV1s, coordinating across multiple partners and network backbone teams, holding them accountable to root/proximate cause analysis and SLAs. •Contribute to capacity planning: projecting GPU demand, managing cluster expansion across GPU generations, and coordinating multi-cluster strategy. Qualifications: •Strong Linux systems fundamentals — networking (TCP/IP, InfiniBand, RDMA, MTU/MSS/PMTUD), storage (NFS, VAST, WEKA, POSIX semantics), kernel internals (cgroups, namespaces, eBPF, sysctls). •Applied hands-on capability in Kubernetes and cloud-native infrastructure — pod lifecycle, CNI plugins (Cilium preferred), StatefulSets, Helm, ArgoCD, or equivalent GitOps tooling. •Debugging instinct: ability to form hypotheses quickly, design controlled experiments, and root cause complex multi-system failures under pressure. You enjoy finding the hard bugs. •Proficiency in Python and Bash for automation and tooling. Go, Rust, or C/C++ a plus. •Excellent communication — you can write a crisp incident summary for researchers, a technical escalation to a vendor CTO, and a system design doc for teammates, all in the same day. •Bonus: experience with distributed AI training infrastructure (NCCL, PyTorch DDP, multi-node job debugging, checkpoint/restart patterns, container environments for large-scale training). Compensation: •$241,000 - $331,000 / year •Provides a generous employer match on employee 401(k) contributions to support planning for the future Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!