Staff Software Engineer, Compute & Storage, Autonomy
Rivian - Palo Alto, CA
Hiring: Staff Software Engineer, Compute & Storage, Autonomy Company: Rivian Location: Palo Alto, CA Job Posted Time: 2026-09-17 05:28:30 Employment Type: Full-time / Hybrid Target Skills & Keywords : AWS, C++, CDK, CloudFormation, CloudWatch, Data Pipeline, Datadog, Event-Driven, Flink, Grafana, Infrastructure as Code, Kafka, Kinesis, Kubernetes, Linux, Prometheus, Python, Rust, SQS, Spark, Systems Design, Terraform About the job Experience: •8+ years of software engineering experience building and operating distributed systems in production. •5+ years owning a distributed compute, batch execution, or scheduling platform used by other engineering teams: building the scheduler, orchestration layer, or execution engine itself, not just jobs that ran on one. •5+ years with Kubernetes as an execution substrate: scheduling, resource management, custom controllers or operators, autoscaling, and failure modes at scale. •5+ years with large-scale object storage: data layout, partitioning, lifecycle and tiering, caching, and the performance and cost tradeoffs among them. •3+ years with a distributed processing or training framework (Ray, Spark, Flink, or Dask). Required Skills: •Own the architecture and roadmap for Autonomy's distributed compute platform: job scheduling, quota and fair-share across teams, autoscaling, and spot and preemption strategy across multiple clouds. •Build scalable tools and APIs that turn high-level job requests into executed work, making large-scale computation accessible to engineers who aren't infrastructure specialists. •Read the jobs other teams run, profile them, and find inefficiencies and bottlenecks. Fix them directly, or give the team the tooling to see them. •Treat cluster efficiency as a primary metric: eliminate GPU fragmentation, right-size quota, and reclaim capacity stranded on partially filled nodes. •Own the storage architecture for autonomy data at petabyte scale, including layout, partitioning, tiering across hot, warm and cold, lifecycle policy, and the caching and prefetch layers that keep training and simulation jobs from starving on I/O. •Own the multi-cloud compute and data path as training extends beyond a single provider, including replication strategy, consistency, and cross-provider egress economics. •Drive throughput and turnaround time for the heaviest workloads: training data loading, large-scale log replay, and batch resimulation running tens of thousands of concurrent jobs. •Own the queue-based and event-driven infrastructure behind job submission and autoscaling, and keep it stable as queue depth moves by orders of magnitude. Qualifications: •Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience. •2+ years carrying an infrastructure cost target and meeting it. •Demonstrated capacity to turn ambiguous, high-level requirements into a detailed system design and drive it to completion unprompted. •Technical influence beyond your own commits: designs others build on, standards others adopt, engineers who improved from working with you. Compensation: •$206,500 - $258,100 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!