Senior Software Engineer, Robinhood Command Center
Robinhood - New York, NY
Hiring: Senior Software Engineer, Robinhood Command Center Company: Robinhood Location: New York, NY Job Posted Time: 2026-09-10 13:39:17 Target Skills & Keywords : Grafana, OpenTelemetry, Prometheus, Spark About the job Experience: •5+ years of software engineering experience, including significant experience operating production systems •2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations Required Skills: •Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure •Partner closely across many different types of engineers to raise the bar for operational excellence and incident response •Lead incident mitigation efforts by coordinating service owners, facilitating time-sensitive decisions like rollbacks, traffic shifts, and maintaining a clear source of truth during active incidents •Develop and maintain incident management processes and procedures to ensure timely resolution and minimize customer impact •Own incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys (CUJs), availability, and business-impact metrics •Own and evolve incident response tooling and processes, including education, adoption, and measurement of MTTD/MTTR improvements •Drive post-incident governance and learning, defining standards for postmortems, SEV reviews, and follow-up tracking to ensure durable reliability improvements •Design and implement next-generation failure mitigation strategies that avoid full-region or full-datacenter failovers Qualifications: •Hands-on experience serving in incident leadership roles (e.g., IMOC, incident commander, primary oncall) •Strong communication and cross-functional collaboration skills, especially during high-severity incidents •Deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design •Operational familiarity with modern observability stacks (e.g., OpenTelemetry, Prometheus, Grafana) •Demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact Compensation: •$196,000 - $230,000 / year •Performance driven compensation with multipliers for outsized impact, bonus programs, equity ownership, and 401(k) matching Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!