Senior Platform Engineer - Frisco

McAfee - Frisco, TX

Hiring: Senior Platform Engineer - Frisco Company: McAfee Location: Frisco, TX Job Posted Time: 2026-09-12 12:18:13 Employment Type: Hybrid Target Skills & Keywords : AWS, Aurora, CI/CD, EKS, Event-Driven, FastAPI, GCP, IAM, Infrastructure as Code, Kafka, Kubernetes, LLM, LangChain, LlamaIndex, MLOps, Machine Learning, Microservices, OpenSearch, Pinecone, Python, RAG, Terraform, VPC, Weaviate, gRPC, pgvector About the job Experience: •10+ years of experience in platform engineering, with hands-on AI/ML or GenAI platform experience. Required Skills: •Design, build, and scale enterprise-grade Generative AI platforms supporting LLM applications, AI agents, RAG architectures, and multi-model routing. •Architect and implement secure, scalable AI infrastructure leveraging cloud-native technologies (AWS, GCP, Kubernetes, GKE/EKS). •Enable self-service AI capabilities for engineering teams through standardized platform services, APIs, and Backstage templates/plugins. •Build and operate Retrieval-Augmented Generation (RAG) infrastructure, including embedding pipelines and vector stores (OpenSearch, Aurora pgvector). •Develop and manage enterprise AI gateway capabilities, including model routing, rate limiting, token tracking, and policy enforcement. •Integrate GenAI services into CI/CD pipelines and platform workflows to enable seamless deployment and lifecycle management. •Build observability platforms for GenAI systems, tracking token usage, latency, response quality, failure rates, throughput, and cost visibility. •Own lifecycle management of Kubernetes-based AI platforms including upgrades, patching, scaling. Qualifications: •Applied hands-on capability in at least one LLM ecosystem (AWS Bedrock, OpenAI, Anthropic). •Strong Kubernetes experience (EKS/GKE), including GPU scheduling, autoscaling, and multi-tenant isolation. •Strong programming expertise in Python and Go; experience building services using FastAPI and gRPC. •Deep expertise in AWS (IAM, VPC, KMS) and Infrastructure as Code (Terraform). •In-depth knowledge of distributed systems and event streaming (Apache Kafka). •Expertise in CI/CD automation and platform engineering best practices. •Exposure to LLMOps / MLOps tooling for model lifecycle management, evaluation, and versioning. •Operational familiarity with AI cost optimization strategies (token efficiency, caching, adaptive routing). •Exposure to AI model evaluation frameworks (quality scoring, hallucination detection, benchmarking). Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!