Machine Learning Engineer | Python | Pytorch | Distributed Training | Optimisation | GPU | Hybrid, San Jose, CA

Enigma - San Jose, CA

Hiring: Machine Learning Engineer | Python | Pytorch | Distributed Training | Optimisation | GPU | Hybrid, San Jose, CA Company: Enigma Location: San Jose, CA Job Posted Time: 2026-09-15 20:05:06 Employment Type: Hybrid Target Skills & Keywords : CI/CD, Data Pipeline, Deep Learning, Load Balancing, Machine Learning, Milvus, ONNX, Parquet, Pinecone, PyTorch, Python, SQL, TensorFlow, TensorRT, Triton, pgvector, vLLM About the job Experience: •5 years in ML/AI engineering roles owning training and/or serving in production at scale. Required Skills: •Productize and optimize models from Research into reliable, performant, and cost-efficient services with clear SLOs (latency, availability, cost). •Scale training across nodes/GPUs (DDP/FSDP/ZeRO, pipeline/tensor parallelism) and own throughput/time-to-train using profiling and optimization. •Implement model-efficiency techniques (quantization, distillation, pruning, KV-cache, Flash Attention) for training and inference without materially degrading quality. •Build and maintain model-serving systems (vLLM/Triton/TGI/ONNX/TensorRT/AITemplate) with batching, streaming, caching, and memory management. •Integrate with vector/feature stores and data pipelines (FAISS/Milvus/Pinecone/pgvector; Parquet/Delta) as needed for production. •Define and track performance and cost KPIs; run continuous improvement loops and capacity planning. •Partner with ML Ops on CI/CD, telemetry/observability, model registries; partner with Scientists on reproducible handoffs and evaluations. •Educational Qualifications: •Bachelors in computer science, Electrical/Computer Engineering, or a related field required; Master’s preferred (or equivalent industry experience). •Strong systems/ML engineering with exposure to distributed training and inference optimization. •Industry Experience: •3–5 years in ML/AI engineering roles owning training and/or serving in production at scale. •Demonstrated success delivering high-throughput, low-latency ML services with reliability and cost improvements. •Experience collaborating across Research, Platform/Infra, Data, and Product functions. •Technical Skills: Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!