Site Reliability Engineer, Observability
Ripple - New York, NY
Hiring: Site Reliability Engineer, Observability Company: Ripple Location: New York, NY Job Posted Time: 2026-09-12 12:04:58 Employment Type: Full-time Target Skills & Keywords : AWS, Agile, Azure, Azure DevOps, Bash, GitHub, Infrastructure as Code, Jira, Linux, New Relic, OpenTelemetry, OpsGenie, PagerDuty, PowerShell, Python, R, REST, SOC 2, SQL, SQL Server, Scrum, Terraform About the job Experience: •Background facilitating chaos engineering or game day exercises to build team resilience •Knowledge of VM-hosted SQL Server monitoring and performance optimization •Operational familiarity with FinTech compliance requirements (SOC 2, ISO 27001) and audit evidence collection •Other common names for this role: Senior Site Reliability Engineer, Observability Engineer, Incident Management Engineer Required Skills: •As a Site Reliability Engineer you will be a force multiplier elevating engineering capabilities across observability and incident management. •Empower Ripple's stream-aligned engineering teams to detect, diagnose, and resolve production issues quickly and effectively—helping keep our products highly available, performant, and resilient at scale for customers managing trillions in annual payment volume. •Be part of Ripple's Technical Operations team, coaching teams to build comprehensive monitoring, effective alerting, and mature incident response practices. Qualifications: •5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering with strong focus on observability and production operations •Proven ability to coach and mentor engineering teams with excellent communication and teaching skills across technical and non-technical audiences •Consultative mindset with the ability to influence and guide teams without direct authority •Expert-level hands-on experience with New Relic (APM, Infrastructure Monitoring, Logs, Synthetics, Alerts) and strong proficiency writing NRQL queries for troubleshooting •Proven experience implementing instrumentation in application code (OpenTelemetry, Serilog, or similar frameworks) •Comprehensive expertise in structured logging, metrics collection (RED/USE methods), distributed tracing, and creating effective dashboards and alerts •Expertise defining and implementing SLOs/SLIs and error budgets for reliability management •Demonstrated ability to troubleshoot complex production issues using observability data across distributed systems •Incident Management Expertise (Required) •Applied hands-on capability in incident management platforms ( Compensation: •Flexible work environment (work from home / hybrid options) •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!