AI Cluster Technical Program Manager – Validation, Debug & Agentic AI

AMD - Austin, TX

Hiring: AI Cluster Technical Program Manager – Validation, Debug & Agentic AI Company: AMD Location: Austin, TX Job Posted Time: 2026-09-10 18:02:29 Target Skills & Keywords : Go, HBase, Make, Move, Node.js, TestNG About the job Required Skills: •We are seeking a Technical Program Manager to lead execution of AI cluster engineering programs with deep focus on GPU platforms, rack-level solutions, and AI Cluster validation. This role is responsible for driving end-to-end delivery from GPU + server integration through rack bring-up, scale testing, failure analysis, and system debug closure, ensuring platform readiness for hyperscale and enterprise AI deployments. •This role operates at the intersection of hardware, firmware, networking, and scale-test execution, and requires strong technical depth combined with disciplined program execution. •You are a hands-on TPM who thrives in complex, fast-moving ecosystems, and can connect deep technical details to crisp program plans, executive reporting, and customer outcomes. You are comfortable driving execution in bring-up and EVT/DVT/PVT working closely with engineers to root-cause issues, unblock debug, and make data-driven tradeoffs to keep programs moving. You bring urgency, ownership, and clarity to ambiguous problem spaces and can communicate effectively from lab floor to executive review. •Key Responsibilities •Program Leadership & Execution •Define, plan, and drive program plans for AI infrastructure systems validation and readiness, including server integration, rack bring-up, and cluster-scale deployment readiness. •Create and maintain core PM artifacts: schedules, dependency maps, resource forecasts, risk/issue logs, and program dashboards/status reports. •Identify and drive mitigation plans for issues/risks, including cross-team escalations and corrective actions across multiple engineering areas. •Drive regular execution reviews with engineering teams and provide concise, data-driven updates to senior leadership. •GPU & Platform Execution •Own program execution for GPU-based AI platforms, spanning system bring-up, qualification, scale readiness, and deployment validation across server, rack, and cluster levels. •Drive alignment across GPU, CPU, firmware, BIOS/BMC, and system teams to ensure readiness for scale testing and customer workloads. •Track platform issues, and debug dependencies; ensure risks are clearly documented, owned, and mitigated. •AI Rack / Cluster Validation •Own program planning and execution for multi-node and multi-rack scale testing, including test strategy, scheduling, coverage tracking, and readiness gates. Compensation: •$162,640 •$243,960 Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!