Real-Time Inference Engineering Lead (FTE / Hybrid)
NTT DATA North America - Charlotte, NC
Hiring: Real-Time Inference Engineering Lead (FTE / Hybrid) Company: NTT DATA North America Location: Charlotte, NC Job Posted Time: 2026-09-03 07:43:03 Employment Type: Remote Target Skills & Keywords : API Gateway, AWS, Accessibility, ArgoCD, Azure, CI/CD, Event-Driven, Feature Store, GCP, Git, GitHub Actions, GitLab CI, Helm, Jenkins, Kubernetes, Load Testing, MLflow, OpenShift, Performance Testing, R, REST, RESTful, Service Mesh, Terraform, Triton, Vertex AI, gRPC About the job Experience: •8+ years of software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering experience. •4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real-time data and ML workloads. •4+ years satrong experience with Kubernetes and container platforms in production, including GKE, OpenShift, or comparable environments. Required Skills: •Define the target architecture and engineering standards for real-time predictive model-serving services across cloud and on-premises environments. •Design, build, test, deploy, and operate scalable online inference services that meet latency, throughput, availability, resiliency, and security requirements. •Establish reusable model-serving patterns for synchronous APIs, asynchronous inference, batch-adjacent processing, and event-driven real-time use cases where appropriate. •Build standardized deployment approaches for predictive models, including model packaging, versioning, release promotion, canary deployment, rollback, and retirement. •Design and implement secure API patterns for inference services, including authentication, authorization, traffic management, rate limiting, auditability, and integration with enterprise systems. •Engineer Kubernetes-based serving platforms using GKE, OpenShift, and related container orchestration capabilities. •Implement autoscaling, resource allocation, quota management, capacity planning, and workload-isolation controls for variable inference demand. •Conduct performance engineering, load testing, stress testing, and failure testing to validate service behavior under expected and peak production workloads. Qualifications: •Demonstrated experience leading technical design and engineering delivery for highly available, performance-sensitive production services. •Strong experience with online inference architecture, model-serving frameworks, or predictive-model deployment patterns. •Hands-on experience designing and operating RESTful, gRPC, or event-driven APIs. •In-depth knowledge of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services. •Demonstrated capacity to work effectively with data science, ML engineering, platform engineering, application teams, security, and business stakeholders. •Online inference and low-latency model-serving architecture. •Model deployment, versioning, routing, rollout, rollback, and lifecycle management. •REST APIs, gRPC, API gateways, authentication, authorization, traffic management, and API observability. Compensation: •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!