Staff GPU Inference SDET
Cerebras - Sunnyvale, CA
Hiring: Staff GPU Inference SDET Company: Cerebras Location: Sunnyvale, CA Job Posted Time: 2026-09-10 13:04:11 Target Skills & Keywords : C++, CI/CD, Firmware, Grafana, Kubernetes, LLM, Node.js, Prometheus, PyTorch, Python, User Experience About the job Experience: •8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer. Required Skills: •Design and implement automated test automation frameworks, regression gates, and release qualification pipelines for the complete GPU inference stack—spanning custom API services, model-serving workers, container runtimes, serving engines, driver stacks, and firmware. •Benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs. decode worker performance, continuous batching, prefix caching, KV-cache efficiency, and tensor/expert parallelism. •Build automated workload replay and benchmarking tools to validate GPU performance models. Track critical serving metrics including Time-to-First-Token (TTFT), Inter-Token Latency (ITL), request throughput, tail latency (P99), and capacity efficiency. •Build validation infrastructure to ensure model accuracy, precision stability (FP16/FP8/quantization), determinism, and output correctness across software updates, kernel fusions, and hardware revisions. •Engineer chaos engineering and fault-injection suites to simulate node failures, inter-node network degradation, GPU memory leaks, driver/firmware mismatches, and automated recovery paths for multi-node GPU clusters. •Integrate automated test pipelines with telemetry tools (e.g., Prometheus, Grafana) to turn one-off investigations into repeatable engineering gates and continuous performance monitoring. Qualifications: •Hands-on experience bringing up, provisioning, and validating multi-node GPU clusters (NVIDIA or AMD ecosystem) across public cloud infrastructure or enterprise data center environments. •Comprehensive expertise in LLM serving engines and distributed runtimes, including prefill vs. decode disaggregation, KV-cache management, and dynamic batching. •Expert-level Python programming skills with extensive experience designing custom test automation frameworks, diagnostic tooling, and CI/CD integration. •Strong proficiency with container orchestration tools (e.g., Kubernetes, Slurm, Ray) and high-performance cluster interconnects (e.g., InfiniBand, RoCE, NCCL). •Proven background in root-cause analysis across software/hardware boundaries, stress testing, and node failure simulation in distributed systems. •Direct experience with either AMD (ROCm / HIP) or NVIDIA software stacks. •Operational familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or C++ Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!