Machine Learning Infrastructure Engineer

Character.ai - San Francisco Bay Area

Hiring: Machine Learning Infrastructure Engineer Company: Character.ai Location: San Francisco Bay Area Job Posted Time: 2026-09-15 17:59:13 Employment Type: Full-time Target Skills & Keywords: ML Infrastructure, GPU Clusters, Kubernetes, Cloud Storage, Compute Engine, PyTorch, TensorFlow, JAX, GPU Kernel Development, High-Performance Computing, LLM Training, Distributed Training, Serving Infrastructure Experience: - 4+ years of experience supporting the infrastructure within an ML environment - Experience designing, building, and maintaining training and serving infrastructure for ML research - Experience developing tools used to diagnose ML infrastructure problems and failures - Experience working with GPUs Required Skills: - Infrastructure support for ML research and product teams - Building tooling to diagnose cluster issues and hardware failures - Monitoring deployments, managing experiments, and supporting research operations - Maximizing GPU allocation and utilization for both serving and training - Cloud platforms (e.g., Compute Engine, Kubernetes, Cloud Storage) Qualifications: - 4+ years of experience supporting infrastructure within an ML environment - Proven experience developing diagnostic tools for ML infrastructure problems and failures - Hands-on experience with cloud platforms (Compute Engine, Kubernetes, Cloud Storage) - Experience working with GPUs Nice to Have: - Experience with large GPU clusters and high-performance computing/networking - Experience supporting large language model training - Experience with ML frameworks like PyTorch/TensorFlow/JAX - Experience with GPU kernel development Compensation: Not specified Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!