Principal System Software Architect, AI/GPU Platforms

AMD - Austin, TX

Hiring: Principal System Software Architect, AI/GPU Platforms Company: AMD Location: Austin, TX Job Posted Time: 2026-09-10 18:02:29 Target Skills & Keywords : HBase, Linux, Node.js, SOC About the job Required Skills: •You will join the system software architecture team behind AMD Instinct™ accelerators, the GPUs powering some of the world's largest AI and HPC deployments. This role is focused on next-generation, rack-scale AI platforms in the MI400 class spanning the GPU, the node, and the scale-up/scale-out fabric that binds thousands of accelerators into a single training and inference system. •As a Systems Software Architect, you sit at the intersection of silicon, firmware, driver, runtime, and framework. You define how the software stack exposes and orchestrates the hardware so that AMD's largest customers can extract maximum performance, reliability, and utilization from their infrastructure. •Key Responsibilities •Own the end-to-end system software architecture for one or more MI400-class subsystems for example GPU memory management, scheduling and queuing, RAS and serviceability, virtualization/partitioning (SR-IOV), or the scale-up/scale-out interconnect software model. •Drive architecture across the stack: kernel-mode driver (amdgpu/KFD), user-mode runtime (ROCr/HSA), firmware interfaces, and the ROCm software platform, ensuring the layers compose cleanly and perform. •Partner with silicon and SoC architects during pre-silicon definition to shape hardware/software interfaces, programming models, and register/firmware contracts before tape-out. •Define the software strategy for multi-GPU and rack-scale topologies, including Infinity Fabric / UALink-style interconnect, collective communication (RCCL), memory coherence, and address translation across the platform. •Establish architecture for reliability, availability, and serviceability at scale, error detection, containment, telemetry, recovery, and graceful degradation across large clusters. •Set direction on performance: identify bottlenecks in the launch path, memory subsystem, and communication path, and define the software mechanisms to close them. •Produce architecture specifications, reference designs, and design reviews that align firmware, driver, runtime, and framework teams onto a shared plan. •Act as a technical anchor across AMD and with strategic hyperscale and AI customers — translating their workload requirements into architectural direction and representing AMD in deep technical engagements. •Preferred Experience •Linux Memory Management and Heterogeneous Memory Management (HMM). •GPU / DRM driver development. •Cache coherence and memory consistency protocols. Compensation: •$204,000 •$306,000 Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!