AI Platform Support Engineer (US)
Lightning AI - New York, NY
Hiring: AI Platform Support Engineer (US) Company: Lightning AI Location: New York, NY Job Posted Time: 2026-09-10 08:36:29 Employment Type: Hybrid Target Skills & Keywords : Grafana, Kubeflow, Kubernetes, Linux, MLOps, Machine Learning, Node.js, OpenTelemetry, Prometheus, PyTorch, Python About the job Experience: •Hands on experience operating machine learning workloads in production or research environments •Operational familiarity with GPU infrastructure and orchestration •Understanding of the operational challenges involved in running ML systems at scale •Strong communication skills and ability to work directly with highly technical customers and engineering teams •Comfortable operating in fast moving, highly ambiguous environments Required Skills: •Work Directly With ML Engineers •Partner directly with customer engineering teams running training and inference workloads in production •Help customers diagnose and resolve complex distributed systems and ML infrastructure issues •Act as a technical advisor during high impact incidents and platform degradation events •Translate infrastructure level issues into actionable guidance for ML engineers •Build credibility with customers through strong technical reasoning and clear communication •Debug ML Infrastructure & Distributed Workloads •Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems Qualifications: •Strong software engineering and systems troubleshooting background •Linux systems knowledge, including networking, storage, process management, and performance tuning Compensation: •Competitive benefits and rewards package •Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!