Link Group is seeking an elite Senior Site Reliability Engineer to join the core team responsible for managing global AI compute infrastructure. You will ensure the reliability, performance, and scalability of our bare-metal servers, high-density GPU racks, and the advanced network backbone that supports our AI platform.
Key responsibilities
- Develop sophisticated Python automation to manage the entire lifecycle of our server fleet, from provisioning to decommissioning.
- Architect seamless operational workflows by integrating core systems like JIRA and PagerDuty to reduce incident resolution times.
- Design and implement a world-class observability stack using Prometheus, Grafana, and OpenTelemetry to predict and prevent hardware failures.
- Lead critical incident responses and conduct blameless post-mortems to drive continuous architectural improvements.
- Utilize AI-driven development tools to enhance technical execution and system performance evaluation.
Requirements
- Extensive background in Site Reliability or Production Engineering managing large-scale, mission-critical infrastructure.
- Exceptional Python programming skills with a focus on building scalable, robust operational tools.
- Deep expertise in networking, including advanced topologies, high-bandwidth routing, switching, and BGP.
- Expert-level experience with modern monitoring stacks and telemetry pipelines.
- Proven ability to lead incident response efforts and define clear technical runbooks.
What we offer
- The opportunity to work on cutting-edge AI hardware and infrastructure at a global scale.
- A collaborative environment that encourages the use of advanced AI utilities in daily engineering tasks.
- Professional growth through solving complex, ambiguous technical challenges in a high-impact role.