Senior/Staff Site Reliability Engineer - Data Center

PathAI - Greater Boston

Hiring: Senior/Staff Site Reliability Engineer - Data Center Company: PathAI Location: Greater Boston Job Posted Time: 2026-09-12 10:46:35 Employment Type: Remote Target Skills & Keywords : Ansible, Configuration Management, Datadog, EKS, Grafana, Machine Learning, Prometheus, S3 About the job Experience: •8 years experience working in physical hardware/facilities, networking, automation or other relevant areas. Required Skills: •Advance the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation. •Design, build and operate our data center to support our rapidly growing Machine Learning team. •Build highly-secure on-premises environments handling NIST/ISO standards. •Integrate on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment. •Improve the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure. •Participate in platform on-call rotations and assist with urgent incident response. Qualifications: •You have a BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or closely related technical field. •You have 8 years experience working in physical hardware/facilities, networking, automation or other relevant areas. •You have demonstrated experience with modern datacenter network designs and comfort operating across network layers. •You’ve administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems). •You have demonstrated experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS-Anywhere/ClusterAPI/KVM). •You have strong expertise in storage solutions and optimizing them for high-performance workloads (e.g., Quobyte, S3, FSx, EFS). •You are highly proficient with automation tools; you eliminate toil by automating everything through scripting, configuration management tools (Ansible/RedFish). •You’ve built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus). •You have a proven operational background managing critical production systems, with extensive experience in incident response, infrastructure scaling, and navigating high-growth challenges. •You have the ability to travel to onsite Datacenter location(s) as needed. Compensation: •$146,250 - $225,000 / year •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!