Senior Site Reliability Engineer
TailorCare - United States
Hiring: Senior Site Reliability Engineer Company: TailorCare Location: United States Job Posted Time: 2026-09-16 15:55:45 Employment Type: Remote Target Skills & Keywords : AWS, CI/CD, CloudWatch, Databricks, Datadog, ECS, EKS, Encryption, Fivetran, GitHub Actions, Grafana, HIPAA, IAM, LLM, Lambda, Python, RDS, REST, S3, SOC 2, Salesforce, Snowflake, Terraform, TypeScript, VPC About the job Experience: •5+ years in Software Engineering, SRE, DevOps, or Platform Engineering, with a proven track record of operating at a Senior level (building systems, delivering technical initiatives, and collaborating with peers). Required Skills: •Implement the standardization of our AWS footprint using Terraform. Eliminate manual provisioning, establish reproducible environments, and integrate AI-assisted tooling (e.g., automated PR reviews, intelligent infrastructure operations) to accelerate our workflows. •Build and maintain universal CI/CD pipelines (e.g., GitHub Actions) that remove friction from the SDLC. Track and improve DORA metrics across engineering teams to ensure we are shipping quickly and safely. •Implement observability improvements across AWS and third-party integrations. Monitor and support SLOs, SLIs, and error budgets for key services, ensuring high availability for our web/mobile apps, telephony stack, and data processing. •Act as a bridge between SRE and software/data engineering. Advocate for operations-focused engineering in an empathetic, contributory spirit. You will treat developers as your customers and build self-service tooling so they can ship without filing tickets. •Because our business operates across all US time zones (ET, CT, MT, and PT), you will help stand up the on-call rotation and contribute to shaping sustainable core on-call hours and escalation paths alongside the Director. •Lead production incidents with a calm demeanor, drive blameless post-incident reviews (RCAs), and collaborate with teams to fix systemic issues so they don't happen again. •Help implement and maintain infrastructure controls (IAM, encryption, network segmentation) that align with HIPAA and HITRUST requirements, ensuring a secure environment for patient data. •When something breaks, you fix it and improve the system so it does not happen again. Qualifications: •Deep hands-on AWS expertise (VPC, IAM, ECS/EKS, Lambda, RDS, S3) and production-grade Terraform experience at scale (modules, state management, multi-environment). •Strong programming skills in Python, Go, TypeScript, or similar. You treat infrastructure as software and automate away toil. •Applied hands-on capability in modern observability stacks and a practical understanding of how to implement SLOs, SLIs, and error budgets in a startup environment. •Deep experience maintaining and standardizing CI/CD pipelines and tracking delivery metrics (like DORA). •Ability and willingness to travel up to 10% as needed for onsite meetings, team collaboration, and company events. Compensation: •From Day 1, employees enjoy medical, dental, vision, life, and disability insurance, wellness resources and an employer HSA contribution Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!