Sr. Staff AI Engineer - On-Prem AI Infrastructure & Agentic Systems

SK hynix memory solutions America Inc. - San Jose, CA

Hiring: Sr. Staff AI Engineer - On-Prem AI Infrastructure & Agentic Systems Company: SK hynix memory solutions America Inc. Location: San Jose, CA Job Posted Time: 2026-09-12 13:54:44 Employment Type: Hybrid Target Skills & Keywords : AI, CI/CD, Docker, Embedded Systems, Fine-tuning, Grafana, Helm, Kubernetes, Linux, Machine Learning, Milvus, ONNX, Ollama, Pinecone, Prometheus, Python, Qdrant, RAG, Systems Engineering, TensorRT, Triton, User Experience, vLLM About the job Experience: •2+ years of experience in AI/ML engineering, with hands-on deployment of AI systems on-prem or private cloud. Required Skills: •Design and deploy on-prem AI infrastructure — including GPU clusters, model serving (e.g., vLLM, TGI, Triton), vector DBs (e.g., Milvus, Qdrant, FAISS), and orchestration (Kubernetes, Helm, Docker). •Build and optimize RAG pipelines — including document chunking, retrieval strategies (hybrid, re-ranking), and evaluation of retrieval accuracy and latency. •Develop agentic AI systems — design stateful agents with memory, tool use, and planning capabilities (e.g., using LangGraph, AutoGen, or custom frameworks). •Fine-tune and deploy embedded models — work with LoRA, QLoRA, or full fine-tuning for domain-specific tasks; optimize for edge/on-device inference. •Implement Model Control Protocols (MCP) — ensure model governance, versioning, access control, and monitoring for production AI systems. •Partner cross-functionally with product and engineering teams to integrate AI capabilities into enterprise workflows — especially in storage, QA, or systems engineering contexts. •Automate and monitor AI pipelines — build CI/CD for model deployment, logging, and performance tracking. Qualifications: •Proven experience building agentic AI systems — including state management, tool integration, and multi-step reasoning. •Strong working knowledge of RAG architectures — chunking, retrieval, re-ranking, evaluation metrics. •Operational familiarity with Model Control Protocols (MCP) or similar governance frameworks (model versioning, access control, audit trails). •Proficiency in Python, Linux, Docker/Kubernetes, and vector databases (e.g., Milvus, Qdrant, Pinecone). •Background in systems engineering or QA automation — bonus if you’ve used AI to automate testing or validation. •Operational familiarity with embedded AI or edge inference. •Knowledge of AI observability tools (LangSmith, Weights & Biases, Prometheus/Grafana for AI). •As a Storage company, knowledge of storage area/NVMe is a PLUS. •Bachelor of Science in CS, EE, ME, or other applicable Engineering field. Compensation: •$140,000 - $165,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!