Principal Software Engineer – PyTorch Training Frameworks

AMD - San Jose, CA

Hiring: Principal Software Engineer – PyTorch Training Frameworks Company: AMD Location: San Jose, CA Job Posted Time: 2026-09-02 18:02:29 Target Skills & Keywords : C, C++, Clean Architecture, HBase, LLM, Node.js, PATs, PyTorch, Python About the job Required Skills: •AMD is looking for a •Principal-level PyTorch training framework expert •to help drive performance, scalability, and correctness of large-scale AI training on AMD Instinct™ accelerators. You will work at the intersection of PyTorch internals, distributed training, and hardware-aware optimization, partnering closely with compiler, kernel, driver, and architecture teams to deliver industry-leading training performance and developer experience. •The ideal candidate is deeply hands-on with PyTorch training and thrives on solving complex systems problems (performance, scaling, memory efficiency, distributed communication). You bring strong technical leadership, can influence architecture across teams, and are comfortable driving ambiguity to crisp execution. You communicate clearly with both engineers and stakeholders and can represent AMD credibly in upstream/open-source discussions. •Key Responsibilities •Act as a technical authority for PyTorch training at AMD, setting direction for performance, scalability, and reliability Drive optimization of key PyTorch training workloads (LLMs/foundation models) across single-node and multi-node systems •Improve and debug training performance in areas such as DDP/FSDP, gradient checkpointing, mixed precision, memory planning, and communication/computation overlap •Partner with ROCm compiler/runtime, kernel, and driver teams to resolve performance bottlenecks and correctness issues across the full stack •Contribute to and influence upstream PyTorch (design discussions, code contributions, performance fixes, CI/debug) •Develop and maintain representative training benchmarks, profiling workflows, and performance regression detection for key models •Lead deep-dive investigations of performance regressions and hard correctness issues; drive cross-team resolution to closure •Mentor engineers and raise the bar on framework-quality code, performance engineering practices, and technical rigor •Engage with strategic customers/partners on training enablement, root-cause analysis, and best-practices for AMD platforms •Preferred Experience •Deep experience with PyTorch internals and training systems (Autograd, optimizers, dataloading, compilation paths, runtime behavior) Compensation: •$240,000 •$360,000 Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!