Lead AI Engineer (FM Hosting, LLM Inference)

Capital One - New York, NY

Hiring: Lead AI Engineer (FM Hosting, LLM Inference) Company: Capital One Location: New York, NY Job Posted Time: 2026-09-16 16:57:03 Employment Type: Full-time Target Skills & Keywords: Foundation Model Hosting, LLM Inference, AI/ML Algorithms, Similarity Search, VectorDBs, Guardrails, Model Evaluation, Experimentation, Governance, Observability, AWS Ultraclusters, Huggingface, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, Cloud Platforms, LLM Optimization, Scalability, Latency, Throughput, Cost Optimization Experience: - Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, OR - Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies - At least 4 years of experience programming with Python, Go, Scala, or Java - Preferred: 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (AWS, Google Cloud, Azure, or equivalent private cloud) - Preferred: Experience designing, developing, delivering, and supporting AI services - Preferred: Experience developing AI and ML algorithms or technologies (e.g., LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang - Preferred: Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Required Skills: - Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability - Leverage Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more - Invent and introduce state-of-the-art LLM optimization techniques to improve performance — scalability, cost, latency, throughput — of large scale production AI systems - Contribute to technical vision and long-term roadmap of foundational AI systems - Strong foundation in engineering and mathematics - Expertise in hardware, software, and AI to identify and exploit optimization opportunities - Ability to intuitively understand scientific publications and judiciously apply novel techniques in production - Adaptability and resilience in bringing clarity to big, undefined problems - Strong collaboration with cross-functional teams including engineers, research scientists, technical program managers, and product managers Qualifications: - Bachelor's or Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields - Passion for staying abreast of the latest AI research and AI systems - Deeply technical with a strong foundation in engineering and mathematics - Courage to share new ideas even when unproven - Resilient trailblazer who can forge new paths to achieve business goals when the route is unknown - Capital One will consider sponsoring a new qualified applicant for employment authorization for this position Compensation: - New York, NY: $215,200 - $245,600 for Lead AI Engineer - Cambridge, MA: $197,300 - $225,100 for Lead AI Engineer - McLean, VA: $197,300 - $225,100 for Lead AI Engineer - San Jose, CA: $215,200 - $245,600 for Lead AI Engineer - Eligible to earn performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI) - Comprehensive, competitive, and inclusive set of health, financial and other benefits Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!