AI Platforms Leader Enterprise AI Platforms
Qualcomm - San Diego, CA
Hiring: AI Platforms Leader Enterprise AI Platforms Company: Qualcomm Location: San Diego, CA Job Posted Time: 2026-09-03 07:38:22 Employment Type: Full-time / Hybrid Target Skills & Keywords : AWS, Airflow, Azure, C, C++, CI/CD, Compliance, Docker, EKS, Feature Store, GCP, GitOps, Helm, IaC, Java, Kubernetes, LangChain, MLOps, MLflow, Milvus, Model Registry, OIDC, OpenAPI, Pinecone, PyTorch, Python, RBAC, SAFe, SageMaker, Service Mesh, Terraform, Triton, Vault, vLLM About the job Experience: •15+ years overall engineering/technology experience, including ~10 years building and operating large‑scale platforms (AI/ML, data, or high‑performance computing). •5+ years, across platform/SRE/MLOps/LLMOps, with coaching, hiring, performance management, and clear execution rhythms. •8+ years of Software Engineering or related work experience. •7+ years of Software Engineering or related work experience. •6+ years of Software Engineering or related work experience. Required Skills: •Own the AI Platform strategy & roadmap •Define the multi‑year vision for a multi‑tenant, hybrid (on‑prem + cloud) AI platform, aligned to business needs, developer productivity, and cost efficiency. •Establish clear platform SLAs/SLOs, reliability goals, and security/compliance guardrails. •Run GPU-based compute at scale •Operate and optimize on‑prem GPU clusters (e.g., Kubernetes + GPU operator and/or Slurm), including capacity planning, scheduling, partitioning, NCCL, and high‑throughput storage/networking. •Drive GPU utilization efficiency, right‑sizing, and cost transparency across training and inference workloads. •Deliver MLOps & LLMOps as a product •Provide golden paths for data prep, training/fine‑tuning, model registry, lineage, governance, evaluation, red‑teaming, and safe deployment (batch, online, streaming). Qualifications: •Leadership: Proven experience leading a team of ~10 engineers for 5+ years, across platform/SRE/MLOps/LLMOps, with coaching, hiring, performance management, and clear execution rhythms. •GPU cluster expertise: Hands‑on operations for on‑prem GPU clusters (Kubernetes + GPU operator and/or Slurm), scheduling, capacity planning, performance tuning, and reliability. •MLOps & LLMOps: Strong experience with model lifecycle (data → training → registry → deployment), model/agent evaluation, safety/guardrails, and observability. •Cloud (AWS/GCP/Azure): Deep experience with AI/ML services and managed Kubernetes (EKS/AKS/GKE), networking, security, identity, and cost management. •DevOps/Platform Engineering: CI/CD, GitOps, IaC (Terraform/Bicep/Helm), containerization (Docker), Kubernetes, and secure SDLC practices. •Agentic AI & MCP: Solid understanding of agent orchestration, A2A patterns, tool abstractions, and operating MCP servers in production. •Operational excellence: Demonstrated success running AI or computing clusters with SLOs, on‑call, incident management, and post‑mortems. •Global collaboration: Experience leading a distributed engineering team across time zones. •Master’s or PhD in CS/EE/Math or related field. •Training & Inference stacks: PyTorch, CUDA/cuDNN, Triton Inference Server, vLLM, KServe, Ray, Slurm. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!