Senior Network Engineer – GPU Cluster Networking
AMD - San Jose, CA
Hiring: Senior Network Engineer – GPU Cluster Networking Company: AMD Location: San Jose, CA Job Posted Time: 2026-09-17 10:14:27 Target Skills & Keywords : AI, BGP, Grafana, Kubernetes, Large Language Model, Linux, Performance Testing, Prometheus, Subnetting, Systems Engineering About the job Experience: •Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments. •Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation •Strong hands-on experience with RDMA and RoCEv2 in production environments. •Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet. •In-depth knowledge of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures. Required Skills: •We are seeking a Senior Network Engineer to join the AMD IT System Engineering team. •The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments. •The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems. •Partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance. •You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,. •You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations. •Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters. •Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs. Compensation: •$173,600 - $260,400 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!