GPU Kernel Engineer
Sciforium - San Francisco, CA
Hiring: GPU Kernel Engineer Company: Sciforium Location: San Francisco, CA Job Posted Time: 2026-09-16 23:47:43 Target Skills & Keywords : C++, LLM, Make, PyTorch, Python, TensorRT, TestNG, Triton, vLLM About the job Experience: •5+ years of industry or research experience in GPU kernel development or high-performance computing. Required Skills: •This role is ideal for someone who thrives at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads, and who wants to make meaningful contributions to the efficiency and scalability of our ML platform. •Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas. •Profile and optimize end-to-end performance of ML operations, with a focus on large-scale LLM training and inference. •Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes. •Develop performance models, identify bottlenecks, and deliver kernel-level improvements that significantly accelerate AI workloads. •Partner cross-functionally with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack. •Work closely with hardware vendors (NVIDIA/AMD) and stay current on the latest GPU architecture capabilities and compiler/toolchain improvements. •Contribute to tooling, documentation, benchmarking suites, and testing frameworks to ensure correctness and performance reproducibility. Qualifications: •Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!