Software Engineer, Spark Platform
DoorDash - Seattle, WA
Hiring: Software Engineer, Spark Platform Company: DoorDash Location: Seattle, WA Job Posted Time: 2026-09-09 18:07:18 Employment Type: Remote Target Skills & Keywords : AWS, Clean Architecture, Databricks, Go, HBase, Java, Kubernetes, Make, Move, Node.js, OpenTelemetry, Prometheus, Python, REST, SQL, Scala, Spark, VPC About the job Experience: •2+ years of industry experience operating production distributed systems. Required Skills: •As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. •You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. •You're excited about this opportunity because you will… •Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. •Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. •Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. •Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sustainable as the team and the workload grow. •Partner with senior engineers on shuffle, runtime, and architecture work, and grow into deeper ownership of those areas over time. •We're excited about you because… •B.S., M.S., or PhD in Computer Science or equivalent. •2+ years of industry experience operating production distributed systems. •Experience operating Apache Spark at scale on Amazon EMR, Databricks, or an in-house deployment — with a focus on platform operations (runtime upgrades, cluster lifecycle, shuffle, observability, multi-tenant scheduling) rather than authoring individual Spark jobs. •Hands-on experience operating production systems on Kubernetes — controllers, operators, custom resources, and the failure modes that show up in multi-tenant clusters. •Familiarity with batch or big-data schedulers (YuniKorn, Volcano, Kueue, or equivalent) and/or with the Spark-on-Kubernetes operator. •Familiarity with observability stacks (Prometheus, OpenTelemetry, distributed tracing, structured logging) and with defining SLOs and SLIs that change team behavior. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!