Senior Site Reliability Engineer, Observability

Ripple - New York, NY

Hiring: Senior Site Reliability Engineer, Observability Company: Ripple Location: New York, NY Job Posted Time: 2026-09-12 12:04:58 Employment Type: Full-time Target Skills & Keywords : AWS, Agile, Azure, Azure DevOps, Bash, CI/CD, IaC, Infrastructure as Code, Jira, Linux, New Relic, OpsGenie, PagerDuty, PowerShell, Python, R, REST, SOC 2, SQL, SQL Server, Scrum, Slack, Terraform About the job Experience: •40 years of experience supporting some of the world’s largest and most sophisticated companies, Ripple Treasury integrates a treasury command center into Ripple’s technology stack—giving corporates the ability to move, manage, and optimize liquidity in real-time, across traditional and digital assets, under one expanded umbrella. •7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations. Required Skills: •Spend the majority of your time doing hands-on observability and reliability engineering work: building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems. •Alongside that, you will coach and consult with stream-aligned product teams, helping them build operational maturity over time. •Join Ripple’s Technical Operations team and work across Azure (80%) and AWS (20%) environments supporting infrastructure that is predominantly Windows-based (80%), handling significant payment volume for enterprise treasury customers. •The incident management program you will help build is early-stage—you will be establishing practices, not inheriting a mature playbook. Qualifications: •Proven ability to deliver hands-on engineering work while coaching and mentoring teams—comfortable switching between builder and consultant modes. •Observability & Incident Management Expertise — Required •Expert-level hands-on experience with New Relic (APM, Infrastructure, Logs, Synthetics, Alerts) and strong NRQL proficiency for troubleshooting and analysis. •Comprehensive expertise in structured logging, metrics collection (RED/USE methods), distributed tracing, and designing effective dashboards and alerts. •Expertise defining and implementing SLOs/SLIs and error budgets for reliability management. •Applied hands-on capability in incident management platforms (Incident.IO, PagerDuty, OpsGenie, or similar). •Demonstrated ability to troubleshoot complex production issues using observability data across distributed systems. •Strong Terraform experience: developing and maintaining IaC for cloud infrastructure and monitoring resources; familiarity with IaC governance patterns. •Proficiency with PowerShell scripting (required given the 80% Windows environment). •Strong experience with Azure cloud (App Services, Virtual Machines, Azure SQL, networking, monitoring) and working knowledge of AWS. Compensation: •Flexible work environment (work from home / hybrid options) •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!