Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization

Netskope - Santa Clara, CA

Hiring: Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization Company: Netskope Location: Santa Clara, CA Job Posted Time: 2026-09-16 11:21:48 Target Skills & Keywords : C++, Fine-tuning, LLM, Machine Learning, ONNX, Python, TensorRT, Zero Trust, vLLM About the job Experience: •10+ years of overall industry experience, with 4+ years hands-on in ML/AI (model development, fine-tuning, and inference optimization). •Hands-on with fine-tuning (e.g. LoRA/QLoRA), quantization (GGUF/AWQ/GPTQ), and inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or MLX/CoreML). On-device or edge inference experience is a strong plus. •Strong Python; comfort reaching into C++ for low-level interop is a plus. •Solid grasp of transformer internals and the levers that move real inference performance and cost: KV cache, attention, batching, memory footprint. •Fluency with agentic coding systems and genuine curiosity about agent harnesses like Claude Code, Pi, and Codex, so you should already be building with them, or itching to. Required Skills: •High-impact ownership. You own the model layer of a net-new product that changes the performance and economics of agentic AI. •Cutting-edge, unusual stack. The hard, interesting inference problems live here: quantization, KV-cache and memory management, sparsity, fine-tuning, and hardware acceleration under real-world resource constraints. •Real scale to build against. Netskope’s customer footprint gives you production signals most teams never see, so you deploy, validate, and iterate fast. •What you will be doing •Build and optimize the model inference path: quantization, KV-cache optimization, batching, and latency/memory/throughput tuning on constrained, commodity hardware. •Fine-tune and evaluate models for bounded tasks; build eval harnesses that gate a capability to release on real accuracy, latency, and security relevance. •Design and grow the task execution runtime (bounded sub-agents), pushing toward dynamic task generation and context compaction. •Drive hardware acceleration / sparsity and support for larger models as the platform matures. Qualifications: •MS in Computer Science, Machine Learning, Electrical Engineering, or equivalent technical degree required, with a focus in AI/ML research; PhD in a related field strongly preferred. Compensation: •$124,500 - $272,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!