Senior Technical Duty Officer, Cloud Ops
Box - Redwood City, CA
Hiring: Senior Technical Duty Officer, Cloud Ops Company: Box Location: Redwood City, CA Job Posted Time: 2026-09-16 12:56:25 Target Skills & Keywords : AWS, Azure, CI/CD, Change Management, DNS, GCP, Grafana, Jira, Kubernetes, Linux, Load Balancing, PagerDuty, Prometheus, Python, SaaS, Shell, Slack, Swift, TLS, Terraform About the job Experience: •5+ years in SRE, production operations, reliability engineering, or equivalent high-scale internet/SaaS operations with repeated experience leading or co-leading major production incidents. Required Skills: •Own and direct live-site Critical and Blocker (and other high severity) incidents from identification and escalation through mitigation and recovery. •Triage, refine, and verify the problem and customer-impact statements. Organize the incident bridge, establish clear swim lanes, coordinate SME’s resources, and lead cross-functional Incident Bridges to mitigate issues and restore service quickly. •Improve Incident Platform tooling - Optimize templates, automate repetitive response steps, and develop tooling to reduce operational overhead and time to mitigate. •Partner with SRE and engineering teams to deepen shared knowledge of Box's dependencies, Tier 1 journeys, failure modes. Work together to implement secure, automated workflows that close operability gaps. •Lead daily change reviews of planned changes in Jira, partnering collaboratively with engineering teams to evaluate and minimize change risk. •Provide day to day technical expertise and experience to the organization to address issues in globally diverse, 24x7 environments - influencing from policy and procedural decisions to key architectural and tooling insights to improve Box's Incident, Change, and Problem Management engineering capabilities. •Lead projects to improve tools and processes related that enhance overall site resiliency, service manageability and observability. •Represent GTOC/NOC in problem management, readiness reviews, and reliability in cross functional forums. Qualifications: •We are an AI-first company. This means you approach your work with a growth mindset and find ways to leverage AI to help make faster, smarter decisions that will 10X your impact at Box. •Demonstrated Incident Commander / Technical Duty Officer (or equivalent) skill: calm under uncertainty, clear communication, ability to delegate, and judgment on when to escalate vs. when to dig deeper. •Strong SRE fundamentals: SLIs/SLOs and error budgets (practical use), observability (metrics, logs, traces), golden signals, blameless postmortems, toil reduction, and “you build it, you run it” partnership with product teams. •Proficient in Python for automation and tooling (not just one-off scripts): readable code, APIs, packaging or service-style tools others can run, and comfort reviewing others’ automation. •Solid Linux/Unix troubleshooting; comfort in multi-tier distributed systems (app, data, network, cloud). •Networking literacy sufficient for real incidents (DNS, TLS, load balancing, HTTP, basic routing/firewall concepts). •Excellent written and verbal communication: exec-ready impact statements, precise Slack/bridge facilitation, and documentation that others can trust under stress. •Proven ability to coach and uplift others — training, mentoring, runbook quality, drills. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!