Link Group is seeking a highly skilled Senior Site Reliability Engineer to join the Virtualized Host Platform team. You will play a critical role in ensuring the reliability, performance, and scalability of our foundational compute platform, focusing on Linux internals and virtualization technologies.
Key responsibilities
- Develop an expert-level understanding of the virtualized host stack, including hardware, BIOS/UEFI, Linux kernel, and KVM/QEMU layers.
- Utilize Python and shell scripting to build robust automation and tooling to manage large-scale infrastructure.
- Design and implement advanced observability systems to proactively identify and resolve issues within the kernel, drivers, or hardware.
- Serve as the primary escalation point for complex system-level investigations, collaborating across teams to perform root cause analysis.
- Participate in an on-call rotation to manage service-impacting incidents and drive long-term system resilience.
Requirements
- Expert-level knowledge of Linux internals, system administration, and the Linux kernel.
- Deep practical experience with virtualization technologies, specifically KVM/QEMU.
- Strong diagnostic skills to troubleshoot complex issues across hardware, firmware, and software layers.
- Proficiency in software development with Python and shell scripting for automation.
- Significant experience in a Senior SRE or development role within large-scale, distributed environments.
- Experience with infrastructure-as-code and configuration management tools like SaltStack or Ansible.
What we offer
- Opportunity to work on a mission-critical, global infrastructure platform.
- A collaborative environment that values deep technical expertise and problem-solving.
- Professional growth through tackling complex, large-scale engineering challenges.