Akamai Technologies is seeking a Senior Site Reliability Engineer to join the Virtualization & Host Platforms team. You will play a critical role in optimizing our cloud platform by identifying performance bottlenecks, managing Linux infrastructure, and developing automation tools to ensure high-quality service for billions of users worldwide.
Key responsibilities
- Develop, test, and distribute automation, software, and services to support the Virtualization & Host Platforms team.
- Design and implement observability infrastructure enhancements to proactively identify and resolve performance issues.
- Create specialized tooling using Ansible or profiling techniques to assist in deep-dive performance investigations.
- Collaborate with cross-functional teams, including kernel and hardware engineers, to troubleshoot complex production problems.
- Participate in on-call rotations to guide the restoration and repair of service-impacting incidents.
Requirements
- Expert-level experience in Linux internals, system administration, and a deep understanding of underlying hardware.
- Advanced knowledge of the Linux kernel, OS, and configuration optimization for KVM/QEMU virtualization.
- Proven experience in designing, developing, and deploying software and infrastructure at scale.
- Strong proficiency with configuration management and orchestration tools such as SaltStack, Ansible, Chef, Puppet, or Kubernetes.
- Bachelor’s degree in Computer Engineering, Computer Science, or equivalent professional experience.
What we offer
- The opportunity to work on a globally distributed cloud platform that powers the world's largest digital experiences.
- A collaborative environment that values creative thinking, technical expertise, and continuous learning.
- Full remote work flexibility, allowing you to perform your best work from the location that suits you.
- Comprehensive support for your health, well-being, and professional growth within a leading technology company.