Principal Site Reliability Engineer (Cloud, Observability & Automation)

The Depository Trust & Clearing Corporation (DTCC) - Jersey City, NJ

Hiring: Principal Site Reliability Engineer (Cloud, Observability & Automation) Company: The Depository Trust & Clearing Corporation (DTCC) Location: Jersey City, NJ Job Posted Time: 2026-09-12 12:41:52 Employment Type: Remote Target Skills & Keywords : AWS, Capital Markets, Dynatrace, ECS, Grafana, Java, Linux, Python, Root Cause Analysis, Splunk, Stakeholder Management About the job Experience: •8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines. Required Skills: •Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications. •Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms. •Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs. •Lead major incident response, root cause analysis, and continuous service improvement initiatives. •Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies. •Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle. •Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives. •Identify operational risks and deliver strategic reliability improvements across the technology ecosystem. Qualifications: •Bachelor's degree in Computer Science, Engineering, or equivalent experience. •Strong hands-on experience with AWS and cloud-native architectures. •Proficiency in Python, Java, Go, or similar programming languages. •Strong Linux/Unix systems administration and troubleshooting experience. •Expertise in observability and monitoring platforms including Splunk, Grafana, Dynatrace, and ITSI. •In-depth knowledge of distributed systems, resiliency engineering, performance tuning, automation, and operational excellence. •Excellent communication and stakeholder management skills with the ability to influence technical and business partners. Compensation: •Flexible work environment (work from home / hybrid options) •Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!