Join the highly skilled Site Reliability Engineering team at Akamai Technologies, where we design and manage the infrastructure powering global compute products. As a Senior Site Reliability Engineer in the Virtualization & Host Platforms team, you will work at the forefront of cloud host and server technologies, ensuring our infrastructure operates at peak performance.
Key responsibilities
- Develop, test, and distribute software, services, and tools for host platform management.
- Design and implement enhancements to observability infrastructure to proactively identify and resolve issues.
- Collaborate with support, operations, and engineering teams to investigate and troubleshoot complex production problems.
- Participate in on-call rotations to guide the restoration and repair of service-impacting issues.
- Automate processes and contribute to the continuous improvement of resource utilization.
Requirements
- Advanced experience in Linux internals, system administration, and deep understanding of underlying hardware.
- Strong knowledge of the Linux kernel, OS, and virtualization technologies such as KVM/QEMU.
- Proven experience in a Development or SRE role, specifically with large-scale distributed systems.
- Proficiency in Python and shell scripting for independent solution implementation.
- Experience with infrastructure management tools like SaltStack, Ansible, Chef, or Puppet.
- Bachelor's degree in Computer Engineering, Computer Science, or equivalent.
What we offer
- Opportunity to work on a highly distributed cloud platform that powers the world's digital experiences.
- Engagement with next-generation hardware and cutting-edge virtualization technologies.
- A collaborative environment focused on solving complex, large-scale engineering challenges.
- Support for professional growth and development within a global engineering team.