Link Group is seeking a highly skilled Senior SRE to join our ACDC Platform team, focusing on infrastructure automation and reliability at an enterprise scale. You will play a pivotal role in designing robust automation solutions and managing our infrastructure as code to ensure high performance and scalability.
Key responsibilities
- Architect and build automation solutions using Ansible, creating playbooks and workflows to eliminate operational toil.
- Manage and enhance the Infrastructure as Code lifecycle using tools like Terraform, Ansible, and SaltStack to ensure repeatable and scalable deployments.
- Engineer for ultimate reliability by developing backend services and proactive observability solutions, including monitoring, logging, and alerting.
- Act as a technical leader and mentor, sharing best practices in automation and reliability across the organization.
- Collaborate with cross-functional teams to troubleshoot complex production issues and participate in incident response.
Requirements
- Expert-level proficiency in Ansible with demonstrable experience in developing playbooks and managing complex configurations.
- Deep understanding of Infrastructure as Code principles and practical experience with tools such as Terraform or SaltStack.
- Advanced Linux engineering and administration skills with a proven track record of deploying infrastructure at a massive scale.
- Strong background as a Site Reliability Engineer or Software Engineer working on large-scale, distributed systems.
- Excellent communication skills with the ability to mentor team members and drive SRE best practices.
What we offer
- Opportunity to work on large-scale, enterprise-grade infrastructure projects.
- Collaborative environment focused on technical excellence and innovation.
- Professional growth through mentorship and leadership opportunities within the platform team.