Akamai Technologies is seeking a Senior Site Reliability Engineer to join our critical AI Hardware SRE team. You will be responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure to ensure best-in-class uptime and reliability.
Key responsibilities
- Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
- Integrating automated workflows across disparate ticketing platforms to enhance resolution times for hardware and network issues.
- Designing and implementing telemetry pipelines, custom monitoring dashboards, and AI-based anomaly detection for bare-metal and virtualized environments.
- Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols.
- Partnering with product teams to optimize scalability, performance, and system reliability while defining key performance indicators.
Requirements
- Solid Computer Science foundation demonstrated through formal education or equivalent practical experience in large-scale SRE or Production Engineering roles.
- Proficient tooling and coding ability in Python to build scalable operational tools, API integrations, and automation frameworks.
- Hands-on experience with modern observability stacks and timeseries engines, such as Prometheus, Grafana, OpenTelemetry, and Loki.
- Working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
- Proven ability to own ambiguous technical challenges, coordinate cross-functional teams, and drive comprehensive, blameless post-mortems.
What we offer
- Opportunity to work on the world's most distributed platform from Cloud to Edge.
- Support for your health, well-being, finances, and life beyond work.
- A culture that trusts employees to work in ways that suit them best.