Sr Staff Engineer

Lattice Semiconductor - San Jose, CA

Hiring: Sr Staff Engineer Company: Lattice Semiconductor Location: San Jose, CA Job Posted Time: 2026-09-17 02:43:25 Target Skills & Keywords : C++, Compliance, Docker, FPGA, Fine-tuning, LLM, Linux, Machine Learning, Milvus, ONNX, OpenAPI, Pinecone, PyTorch, Python, R, RAG, TensorFlow, TensorRT, Weaviate, vLLM About the job Experience: •8+ years of experience in AI and machine learning, with at least 3 years of experience working on LLMs, code generation, large-scale neural networks, RAG, or AI-powered automation. •At least 3 years of experience working on LLMs, code generation, large-scale neural networks, RAG, or AI-powered automation. Required Skills: •Train, fine-tune, and evaluate machine learning models, including LLMs, using techniques such as supervised fine-tuning (SFT), LoRA/QLoRA, and RLHF. •Optimize models for local inference through quantization, pruning, and distillation. •Deploy models on-prem or at the edge using frameworks such as PyTorch, TensorRT, ONNX, vLLM, or llama.cpp. •Build and maintain training and inference pipelines for reproducibility and scalability. •Integrate locally deployed models into production systems via APIs and internal services. •Monitor model performance, drift, latency, and resource utilization in production. •Partner cross-functionally with software engineers, infrastructure teams, and domain experts to deliver end-to-end AI solutions. •Ensure models meet security, privacy, and compliance requirements, especially in restricted or offline environments. Qualifications: •Master’s or Ph.D. in computer science, Engineering, or a related field, or equivalent practical experience. •8+ years of experience in AI and machine learning, with at least 3 years of experience working on LLMs, code generation, large-scale neural networks, RAG, or AI-powered automation. •Applied hands-on capability in LLMs (e.g., GPT-OSS, Nemotron, Kimi, Qwen, LLaMA, Mistral, Falcon, or similar open-weight models). •Proficiency in Python and ML frameworks such as PyTorch or TensorFlow. •Expertise in vector databases (FAISS, Weaviate, Chroma, Pinecone, Milvus) and retrieval models •Experience deploying models in local, on-prem, or resource constrained environments. •Practical experience utilizing multi-agent AI systems (LangGraph, CrewAI, AutgoGen, OpenAI Assistants API) for autonomous coding tasks •Hands-on model development, working with business stakeholders to define KPIs and develop and deliver multi-modal (Text and Images) and ensemble models. •Solid understanding of model optimization techniques (quantization, batching, memory optimization). •Operational familiarity with Linux, containers (Docker), and basic cloud/on-prem infrastructure concepts. Compensation: •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!