Staff Software Engineer, Infrastructure
Docker, Inc - United States
Hiring: Staff Software Engineer, Infrastructure Company: Docker, Inc Location: United States Job Posted Time: 2026-09-12 11:36:54 Employment Type: Full-time / Remote Target Skills & Keywords : ArgoCD, CI/CD, Docker, EKS, Envoy, GitHub Actions, GitOps, Grafana, Kubernetes, Linux, OpenTelemetry, Prometheus, R, REST, Terraform About the job Experience: •8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering. Required Skills: •Take ambiguous infrastructure problems and turn them into proposals the org can rally around, then drive them through RFCs and architecture reviews across teams. •Design self-service capabilities and platform APIs (primarily in Go) for onboarding, provisioning, deployment, observability defaults, and day-2 operations, with contracts and docs teams actually use. •Set delivery standards using Terraform, GitOps with Argo CD, progressive rollout, and good testing, including building the continuous-deployment flow we're missing today. •Evolve the multi-tenant EKS foundations toward better reliability, security, scale, and cost: Envoy Gateway ingress, traffic routing, and the multi-region, cross-account connectivity we need. •Improve SLOs, alerting, and incident follow-up on Grafana Cloud so production gets safer and less dependent on heroics. •We judge this work by outcomes the consuming teams feel: how fast they can provision and ship, how much they can do without us, and how reliably it all runs. •We're Actively Investing In AI-assisted And Agentic Workflows To Cut Operational Toil. We Care That They Stay Safe, Auditable, And Human-reviewed. You'll Help Shape Where These Earn Their Place And Where They Don't. Early Targets Include •Alert enrichment and incident context-gathering: assembling the relevant signals, history, and runbook so the on-call engineer starts with context instead of a blank page. Qualifications: •Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience •Strong software engineering in Go or a similar language: design, testing, debugging, review, long-term maintainability. •A track record designing, shipping, and operating cloud services or infrastructure platforms in production. We hire for skill and impact, not years. •Deep expertise in at least one of: Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms, plus solid Linux, networking, and production-ops fundamentals. •Clear written and verbal communication in a remote environment (RFCs, design docs, incident writeups). •EKS and ingress/CNI/service-mesh experience; observability with OpenTelemetry/Prometheus/Grafana; CI/CD and progressive delivery (GitHub Actions, Argo CD, canaries); experience leading migrations or adoption programs across teams. •You don't need every item here. We value deep expertise in one area, strong systems judgment, and curiosity across the rest. •Build context, meet partner teams, ship your first change, shadow on-call. •Own a strategic platform problem with a clear plan and metrics; lead an improvement from design to production. •One Year Outlook Compensation: •Flexible work environment (work from home / hybrid options) •Competitive benefits and rewards package •Time to recharge – Generous PTO, designated quarterly Whaleness Days, and a designated end-of-year Whaleness break Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!