Senior Inference Engineer - AI

Thomson Reuters - Eagan, MN

Hiring: Senior Inference Engineer - AI Company: Thomson Reuters Location: Eagan, MN Job Posted Time: 2026-09-10 10:40:14 Employment Type: Hybrid Target Skills & Keywords : AI, AWS, AutoCAD, Azure, C++, CAD, CI/CD, Cloud Native, Deep Learning, GCP, Kubernetes, LLM, Microservices, OCI, ONNX, OpenSearch, PyTorch, Python, Snowflake, TensorFlow, TensorRT, Vertex AI About the job Experience: •5+ years of relevant experience Required Skills: •As a Senior Inference Engineer, AI •Within Platform Engineering and Enterprise AI Services, an AI Inference Engineer is responsible for productionizing, optimizing, and scaling AI and LLM workloads that power TR’s AI driven products. •This role ensures that our trained models—from classical ML to generative AI—run efficiently across TR’s multi cloud footprint (AWS, Azure, GCP, OCI), meet strict enterprise reliability requirements, and integrate seamlessly with our data backbone (Snowflake, OpenSearch vector search, API managed model routing). •The successful candidate will help build the next generation of TR’s AI infrastructure, working alongside cloud engineering, data engineering, product teams, and AI Services. •Optimize LLMs and ML models for high-performance inference using techniques such as quantization, pruning, distillation, and hardware specific tuning •Deploy and scale inference workloads on GPUs across AWS, Azure, GCP and internal Kubernetes clusters, ensuring predictable performance during peak traffic hours, especially during business hours •Implement routing and failover strategies for OpenAI/Anthropic/Vertex AI traffic •Integrate models into production grade APIs supporting TR products and enterprise workflows. Qualifications: •In-depth knowledge of ML/LLM fundamentals and inference optimization techniques. •Applied hands-on capability in GPU programming (CUDA preferred), inference runtimes (TensorRT, ONNX •Runtime), and deep learning frameworks (PyTorch/TensorFlow) •Proficiency in Python and at least one systems language (C++ strongly preferred for performance critical inference paths) •Operational familiarity with vector search systems (OpenSearch vectors) and retrieval augmented generation pipelines •Knowledge of distributed systems, microservices, CI/CD, and cloud native architecture •New Position: This position is open due to an existing vacancy to support our evolving business needs. •Hybrid Work Model: We’ve adopted a flexible hybrid working environment (2-3 days a week in the office depending on the role) for our office-based roles while delivering a seamless experience that is digitally and physically connected. •Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrow’s challenges and deliver real-world solutions. Our Grow My Way programming and skills-first approach ensures you have the tools and knowledge to grow, lead, and thrive in an AI-enabled future. •Industry Competitive Benefits: We offer comprehensive benefit plans to include flexible vacation, two company-wide Mental Health Days off, access to the Headspace app, retirement savings, tuition reimbursement, employee incentive programs, and resources for mental, physical, and financial wellbeing. Compensation: •Flexible work environment (work from home / hybrid options) •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!