Sr. SWE Datacenter Automation
Zipline - South San Francisco, CA
Hiring: Sr. SWE Datacenter Automation Company: Zipline Location: South San Francisco, CA Job Posted Time: 2026-09-10 10:33:46 Target Skills & Keywords : Ansible, BGP, CI/CD, Firmware, GitOps, Helm, Kubernetes, Node.js, Python, Rust, Terraform About the job Experience: •5+ years of engineering experience with at least 4 years owning production datacenter, virtualization, or infrastructure automation systems. •At least 4 years owning production datacenter, virtualization, or infrastructure automation systems. Required Skills: •Own end‑to‑end lifecycle for datacenter compute and storage: bare‑metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle. •Design, build, and operate automation that reduces manual setup time and increases deployment velocity: PXE/firmware workflows, dynamic inventory, image generation, fleet-wide configuration drift detection, and automated recovery playbooks. •Deliver measurable reliability and scale improvements: set SLIs/SLOs for provisioning time, node commissioning success rate, cluster upgrade success rate, and mean time to recover (MTTR); own meeting those targets. •Lead cross‑functional runbook and incident ownership for infra incidents affecting flight operations or telemetry: on‑call rotation, incident commander for datacenter platform incidents, postmortems and action items. •Instrument and maintain monitoring, alerting, and dashboards for hardware health, hypervisor performance, storage latency, Kubernetes control plane health, and cluster autoscaling behavior. •Implement cost, capacity, and lifecycle management: capacity planning for compute/storage, automated reclamation, firmware/BIOS/hypervisor patch pipelines, and cold‑standby / failover procedures for critical systems. •Execute hands‑on tasks when required: racking and cabling in datacenters, troubleshooting hardware failures, capture forensic logs, and coordinate physical repairs with vendors and field ops. Qualifications: •Deep, hands‑on expertise with bare‑metal provisioning and imaging (PXE/iPXE, IPMI, Redfish), hypervisors (KVM/qemu, ESXi or equivalent), and storage systems (Ceph, NVMeoF, SAN) at scale. •Proven Kubernetes operations experience: cluster provisioning, upgrades, control‑plane HA, kubeadm/cluster API or equivalent, CNI and CSI troubleshooting, and workload scheduling at multi‑cluster scale. •Production‑grade automation and coding skills in one or more languages (Python, Go, or Rust) and experience with CI/CD pipelines, Terraform/Ansible/Helm, and GitOps practices. •Strong networking fundamentals: VLANs, BGP/EVPN at leaf/spine, LACP, routing, and network troubleshooting for cluster networking and storage fabrics. •On‑call and incident experience: you have owned postmortems, SLIs/SLOs, and driven reliability improvements under operational pressure. •Physical datacenter readiness: able to work on‑site in South San Francisco HQ with regular in‑office cadence, plus occasional travel to partner datacenters or field sites and hands‑on rack/cable/repair work when required. •Security and safety mindset: experience operating in regulated or safety‑sensitive environments, following change control and audit processes. •Clear communication and cross‑team ownership: you will partner with flight software, field ops, hardware, and SRE teams and must translate operational needs into automated, testable systems. •What Else You Need To Know •This role is based in South San Francisco with an expectation of regular on‑site presence and participation in on‑call rotations and occasional travel to datacenter or field locations. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!