Service Operations Director - Evinova

Evinova - Gaithersburg, MD

Hiring: Service Operations Director - Evinova Company: Evinova Location: Gaithersburg, MD Job Posted Time: 2026-09-09 17:57:01 Employment Type: Full-time Target Skills & Keywords: Incident Management, Technical Program Management, Platform/Cloud Operations, ITIL, On-Call Tooling, Observability & Monitoring, AWS, SaaS Architecture, Post-Incident Reviews, ITSM, SLO/SLI, Error Budgets, Stakeholder Communications, Documentation Publishing, Health Tech Experience: - 8+ years of experience in incident management, technical program management, or platform/cloud operations in an engineering organization - Proven ability to lead cross-functional teams calmly and decisively under pressure - Experience communicating with senior leadership during live incidents - Experience running post-incident reviews and building a culture of blameless continuous improvement - Highly preferred: background in health tech, regulated software environments, or high-availability platforms - Highly preferred: experience managing distributed, multi-timezone on-call rotations and global handoff protocols - Highly preferred: experience building or owning documentation publishing workflows across engineering organizations Required Skills: - Own all production incidents end to end as the single accountable coordinator, leading War Room response sessions - Coordinate internal incidents including deployment failures, pre-production degradation, staging outages, and P1/P2 tooling or environment issues - Apply and refine tiered incident response plans based on severity and customer impact - Authorize emergency change actions (ECA) during incidents - Chair blameless post-incident reviews with clear action items and follow-through - Develop and maintain incident response playbooks, runbooks, and escalation paths - Own on-call processes, tool configuration, schedules, and global handoff protocols - Track and report incident KPIs (MTTD, MTTR, repeat incidents) and present trends to leadership - Own stakeholder and executive communications during active incidents - Own customer-facing incident communications including status page updates and post-incident summaries - Own the process for publishing product documentation, user guides, and release notes - Hands-on experience with on-call tooling (PagerDuty, Splunk On-Call, or equivalent) - Hands-on experience with observability and monitoring platforms (Datadog, Splunk, CloudWatch, Grafana, or equivalent) - Strong understanding of cloud infrastructure (AWS preferred) and modern SaaS architecture - Familiarity with ITIL or similar incident and change management frameworks - Highly preferred: experience with ITSM platforms (ServiceNow, Jira Service Management, or equivalent) - Highly preferred: familiarity with SLO/SLI and error budget frameworks - Highly preferred: experience with software engineering practices to help close down issues preventing engineers from working Qualifications: - Bachelor's Degree in a related field - Excellent written and verbal communication skills - Ability to thrive in a fast-paced, high-impact scale-up environment - Flexibility to support global customers across multiple time zones, with occasional morning or evening commitments - Expected to work from the office three days per week Compensation: - Annual base pay range: $166,000 - $220,000 USD (varies based on market location, job-related knowledge, skills, and experience) - Short-term incentive bonus opportunity - Eligibility to participate in equity-based long-term incentive program - Benefits include a qualified retirement program (401(k) plan); paid vacation and holidays; paid leaves; and health benefits including medical, prescription drug, dental, and vision coverage Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!