Senior Staff Machine Learning Software Engineer

San Diego Stealth Startup - San Diego, CA

Hiring: Senior Staff Machine Learning Software Engineer Company: San Diego Stealth Startup Location: San Diego, CA Job Posted Time: 2026-09-16 03:28:57 Employment Type: Full-time Target Skills & Keywords : C++, Embedded Systems, Python, Rust, TensorRT About the job Experience: •6+ years), MS (10+ years) or BS/BA (12+ years) of experience in life sciences or technology. Required Skills: •If your best work is making inference faster, smaller, more predictable, and easier to ship, this role is likely a good match. •Turning research prototypes into production inference components with explicit latency, throughput, memory, and accuracy budgets •Optimizing the execution path: tensor layout, host/device transfers, batching strategy, kernel launch overhead, mixed precision, quantization, and memory reuse •Writing or tuning Rust, C++, and CUDA where framework-level optimization is not enough, then validating the improvement with profiler output and release-facing tests •Building inference-adjacent evaluation machinery: calibration checks, confidence behavior, regression detection, dataset slices, and failure-mode reporting tied to product metrics •Maintaining the deployment contract: model artifacts, runtime integration, versioning, reproducibility, and performance gates that block unsafe changes •Algorithm research and novel model design live on a separate track. You will collaborate with that team, translate prototypes into production constraints, and surface shipping risks early when a design needs to change. Qualifications: •PhD (6+ years), MS (10+ years) or BS/BA (12+ years) of experience in life sciences or technology. •Required proficiency in demonstrated leadership or ownership with 2 of the 5 areas referenced below successfully: •Shipped constrained inference. You have personally moved a model or learned component from prototype to deployed runtime with a real latency, throughput, memory, or power budget. You can name the target, the bottleneck, and the change that closed the gap. •Rust/C++ at shipping depth. You have written production code in Rust or modern C++ where correctness, latency, memory layout, and ownership boundaries mattered. You can reason about the runtime behavior of the code you ship, not just its API surface. •CUDA and accelerator-aware execution. You are comfortable below Python: custom CUDA extensions or kernels, host/device memory movement, launch overhead, profiler traces, and the practical tradeoffs between framework convenience and a purpose-built implementation. •Performance-native judgment. You reason in wall-clock time, memory movement, launch overhead, bandwidth, numerical precision, and error budgets without needing those constraints added late in review. •Production engineering discipline. You define typed interfaces, deterministic behavior, reproducible artifacts, meaningful tests, and clean handoffs with upstream research code. •Rust at shipping depth, especially FFI boundaries, pyo3 / maturin, async runtimes, or performance-sensitive service code •Inference on constrained local hardware, embedded systems, edge devices, or budget-bound accelerator deployments •Quantization, mixed precision, model compression, or kernel fusion that shipped beyond a benchmark notebook Compensation: •$202,000 - $215,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!