Link Group is seeking a seasoned Senior Site Reliability Engineer to join the team responsible for the backbone of global AI/ML services. You will act as a guardian of a massive, distributed AI compute platform, ensuring that GPU-powered infrastructure and AI models are fundamentally reliable, observable, and built to scale.
Key responsibilities
- Design and implement comprehensive observability systems using Prometheus and Grafana, defining and tracking SLOs and SLIs to ensure service reliability.
- Automate manual processes using Python or Go to create self-healing systems and accelerate incident response.
- Lead incident management efforts, participate in blameless on-call rotations, and conduct post-mortems to drive lasting architectural improvements.
- Develop and maintain robust CI/CD pipelines with automated safety checks and rollback capabilities to support rapid innovation.
- Partner with product engineering teams to provide guidance on reliability, shaping architectures for operationally sound AI services.
Requirements
- Proven track record in Site Reliability, Platform, or Infrastructure Engineering managing complex, large-scale distributed systems.
- Deep, practical experience managing large-scale containerized environments using Kubernetes under heavy load.
- Strong programming skills with the ability to write clean, scalable automation scripts and infrastructure-as-code using Python, Go, or Terraform.
- Direct experience or strong interest in AI/ML infrastructure challenges, including model serving pipelines and GPU-accelerated workloads.
- Excellent problem-solving skills with a focus on long-term prevention and a collaborative approach to mentoring team members.
What we offer
- Opportunity to work on cutting-edge AI/ML infrastructure at a massive scale.
- A culture that prioritizes blameless post-mortems and continuous technical improvement.
- Full remote work environment with a focus on autonomy and professional growth.