Member of Technical Staff - GPU Performance Engineer

Liquid AI - San Francisco, CA

Hiring: Member of Technical Staff - GPU Performance Engineer Company: Liquid AI Location: San Francisco, CA Job Posted Time: 2026-09-16 11:30:52 Target Skills & Keywords : C, C++, PyTorch, Triton About the job Experience: •Authored custom CUDA kernels (not only calling cuDNN/cuBLAS) •In-depth knowledge of GPU architecture and performance: memory hierarchy, warps, shared memory/register pressure, bandwidth vs compute limits •Proficiency with low-level profiling (Nsight Systems/Compute) and performance methodology Required Skills: •While San Francisco and Boston are preferred, we are open to other locations. •Works profiler-first: You use tools like Nsight Systems / Nsight Compute to find bottlenecks, validate hypotheses, and iterate until improvements show up in end-to-end benchmarks. •Bridges theory and practice: You can translate ideas from papers into implementations that are robust, testable, and performant. •Executes independently: Given an ambiguous bottleneck, you can drive from profiling to kernel/integration changes to benchmarked results to maintained ownership. •Cares about the details: Memory hierarchy, occupancy, launch configs, tensor core utilization, bandwidth vs compute limits. •Write high-performance GPU kernels for our novel model architectures •Integrate kernels into PyTorch pipelines (custom ops, extensions, dispatch, benchmarking) •Profile and optimize training and inference workflows to eliminate bottlenecks Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!