Senior Software Engineer — Infra Agent Systems
Together AI - San Francisco, CA
Hiring: Senior Software Engineer — Infra Agent Systems Company: Together AI Location: San Francisco, CA Job Posted Time: 2026-09-11 10:59:36 Employment Type: Full-time / Remote Target Skills & Keywords : API Design, ArgoCD, Event-Driven, Fine-tuning, GitOps, Grafana, Kafka, Kubernetes, LLM, NATS, Prometheus, Python, RAG, Reinforcement Learning, Rust, Salesforce, Slack, Systems Design, TypeScript, Zoom About the job Experience: •5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Required Skills: •Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. •We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. •You’ll Work Across Two Areas •Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. •Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. •We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. •This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation •, solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. Qualifications: •Strong systems design skills and experience owning significant systems from design through production. •AI agent systems, orchestration, tool use, evaluation, or grounding •Knowledge graphs or graph data modeling •Search, retrieval, ranking, RAG, or semantic search systems •Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems. •Comfortable working across languages such as Go, TypeScript, Python, or Rust. •GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers •Event-driven systems and messaging platforms such as NATS or Kafka •Observability platforms such as Prometheus and Grafana •Building evaluation frameworks or improving the quality and reliability of LLM-powered systems Compensation: •$250,000 - $300,000 / year •Flexible work environment (work from home / hybrid options) •We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!