Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer

Electronic Arts (EA) - Redwood City, CA

Hiring: Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer Company: Electronic Arts (EA) Location: Redwood City, CA Job Posted Time: 2026-09-16 13:54:10 Employment Type: Full-time / Hybrid Target Skills & Keywords : AI, AWS, AutoCAD, Bash, CAD, EC2, EKS, Grafana, IAM, Infrastructure as Code, Kubernetes, Node.js, PowerShell, Prometheus, Python, S3, Terraform, VPC About the job Experience: •8+ years of experience operating production infrastructure, with deep, current, hands-on AWS depth — EC2 GPU fleets, EKS, IAM and cross-account security, VPC and networking, S3 and FSx Required Skills: •Own GPU fleet operations across our AWS estate. •Build the scheduling layer from zero. •Diagnose GPU and node failures fast and completely and drive hardware evidence and replacement through AWS support and capacity-block channels. •Run researcher support as a first-class product including holding office hours, owning the support channel, and driving the recurring causes out of existence with self-service tooling, preflight checks, and documentation •Instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics. •Partner with our external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation. •Author runbooks, decision records, and onboarding docs. Qualifications: •Electronic Arts creates next-level entertainment experiences that inspire players and fans around the world. Here, everyone is part of the story. Part of a community that connects across the globe. A place where creativity thrives, new perspectives are invited, and ideas matter. A team where everyone makes play happen. •« Pour visualiser la description de poste en français, veuillez sélectionner le français dans le menu déroulant au haut de la page. » •As Lead Infrastructure Engineer, you will own the GPU fleet our researchers train on including capacity, scheduling, diagnostics, and support. You will set technical direction for GPU operations and infrastructure architecture. You will additionally lead Infrastructure as Code setup, granting permissions, and debugging infrastructure problems. •This is a hybrid role, working three days per week in Redwood City, Montreal, or Vancouver. •Report to the Head of Data and Infrastructure. •Expertise in scripting and automation with Python, PowerShell, bash, or equivalent •Expertise in infrastructure as code (Terraform or equivalent) •Operational familiarity with a GPU scheduling or orchestration layer (like Slurm, Kubernetes with Kueue or Volcano, Ray, dStack or SkyPilot) •Observability practice including Grafana, Prometheus, or equivalent Compensation: •$169,500 - $242,600 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!