Senior Site Reliability Engineer
Akamai Technologies - Cambridge, MA
Hiring: Senior Site Reliability Engineer Company: Akamai Technologies Location: Cambridge, MA Job Posted Time: 2026-09-16 12:56:53 Employment Type: On-site Target Skills & Keywords : BGP, Grafana, LLM, OpenTelemetry, PagerDuty, Prometheus, Python, Slack About the job Experience: •5 years of relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent Required Skills: •Do you enjoy collaborating with teams to solve complex challenges? •Do you enjoy solving large scale distributed content delivery challenges? •Responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. •Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning. •Integrating automated workflows across disconnected corporate ticketing systems to optimize time-to-mitigate metrics for hardware and network break-fix events. •Leveraging advanced AI utilities and LLM-assisted development paradigms where appropriate to accelerate technical execution, script authorship, and system analysis •Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments. •Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments. Compensation: •$121,400 - $218,600 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!