Director - Application Site-Reliability Engineering

Caris Life Sciences - Irving, TX

Hiring: Director - Application Site-Reliability Engineering Company: Caris Life Sciences Location: Irving, TX Job Posted Time: 2026-09-16 14:20:59 Employment Type: Hybrid Target Skills & Keywords : CI/CD, HIPAA About the job Experience: •10+ years of professional experience in SRE, DevOps, platform engineering, or production operations. •4+ years of direct experience in a people-management or team-lead role within an SRE or production-operations function. Required Skills: •The infrastructure organization owns the platform and observability runtime; this role owns application-layer reliability on top of it. The operating model is automation-first: recurring manual work is engineered away rather than staffed, and operational and compliance evidence is produced by pipelines rather than assembled by hand. •Own the production-support model for the clinical application portfolio: the on-call rotation, escalation procedures, and incident-response playbooks. •Establish and audit production-access and segregation-of-duties controls with engineering, information-security, quality, and infrastructure partners; keep the evidence audit-ready for SOX ITGCs and applicable FDA requirements, including access grants, role changes, and privileged-action logs. •Define clinical SLOs, error budgets, and availability targets with product and engineering leadership; track attainment and operational-health metrics such as mean time to detect, mean time to recover, and on-call burden; intervene while an error budget is burning, not after it is spent. •Lead incident response for high-severity production events as incident commander or senior technical responder; coordinate cross-functional teams, run post-incident reviews, and drive systemic remediation to closure. •Develop the runbook library and approved operational automation; ensure engineers can execute standard interventions safely within documented procedures. •Strategically drive automation-first operating model: convert recurring manual interventions into reviewed automation, measure and reduce toil, and favor self-service tooling over ticket-driven request work. •Coordinate production deployments with engineering teams, verify deployment health, and own rollback decisions. Qualifications: •Bachelor's degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience. •Hands-on experience leading incident response for Tier 1 or business-critical production systems, including serving as incident commander or senior technical responder. •A record of applying AI-assisted practice to operations or engineering work, personally or through a team. •Domain experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software. •Direct experience supporting SOX ITGC audit cycles or CAP/CLIA laboratory inspections, including evidence gathering for access-control, change-management, and monitoring controls. •Solid functional working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms. •Demonstrated capacity to operate as a player-coach, contributing directly to technical work while building and leading a team. •Track record of reducing operational toil through automation programs in an SRE or production-operations organization. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!