AI Infrastructure Engineer
Percepta - New York, NY
Hiring: AI Infrastructure Engineer Company: Percepta Location: New York, NY Job Posted Time: 2026-09-16 20:57:31 Employment Type: Contract Target Skills & Keywords : AWS, Azure, Bash, CI/CD, Docker, Embedded Systems, GCP, GitHub Actions, GitLab CI, GitOps, Grafana, HIPAA, IAM, IaC, Kubernetes, MLOps, Prometheus, Python, REST, SOC 2, SaaS, Terraform About the job Experience: •Deep experience with at least 1 major cloud provider (AWS, GCP, or Azure): networking, IAM, cost management, the operational realities of production workloads •Solid Docker and Kubernetes experience in production. We run managed clusters across all 3 major clouds; this is a core part of the role •Scripting proficiency in Python, Bash, or similar •High agency: you don't wait for a ticket to fix what's broken, but you communicate, collaborate, and bring the team along •Genuine curiosity about AI systems, not just the infrastructure running them. You want to understand what you're operating Required Skills: •The infrastructure patterns for the agentic systems of the future don't exist yet. You'll help define them. •You're deploying autonomous systems. The infrastructure contract changes when your workloads have agency. •Observability means understanding why an agent made a decision, not just whether a pod is healthy. •The gap between research and production is real here. Our teams move optimization algorithms and AI systems from research environments into production, and you'll be part of that handoff. MLOps experience isn't required, but you'll be closer to that boundary than most infra roles. •Small team. Real ownership. You're making foundational decisions, not inheriting someone else's. •Define infrastructure patterns for multi-agent systems that need to be observable, controllable, and recoverable in ways traditional apps don't require •Own and evolve our IaC stack: Terraform and Kubernetes across AWS, GCP, and Azure •Build observability primitives for agentic workflows, tracing agent decisions and execution paths, not just service latency and pod health Qualifications: •Multi-region and multi-cloud experience across 2+ providers •Operational familiarity with GitOps patterns and progressive delivery •Operational familiarity with the Grafana stack (Prometheus, Grafana, Loki) or equivalent •Background supporting ML or research workflows moving to production: model deployment, pipeline orchestration, or similar •You've thought about what observability means for non-deterministic systems and have opinions about it •The infrastructure patterns for autonomous AI systems are still being written. If you want to be one of the people writing them, let's talk. •We’re working against an incredibly ambitious mission. It won’t be easy, but it will likely be the most fulfilling work of your career. If this excites you, let's chat, even if you don't meet all of the qualifications above. •We have the unique privilege of taking on the most ambitious problems and we should chase them with optimism, responsibility, and genuine belief that we can make it happen. We have to embrace the hard things when no one else will. •Everyone is an engineer and the job of an engineer is to deliver outcomes, not outputs. Everything we do—the products we build, the partnerships we launch, the strategy we set—exists to make our customers successful. Delivery is the strategy. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!