Spin is seeking a highly experienced Senior Site Reliability Engineer to lead the enhancement and maintenance of our IT infrastructure and applications. In this role, you will focus on advanced system stability, efficiency, and scalability, driving strategic initiatives while mentoring junior team members to ensure the highest standards of operational excellence.
Key responsibilities
- Design and maintain sophisticated monitoring solutions to ensure optimal health and performance of infrastructure and applications.
- Lead complex incident response activities, diagnose critical reliability issues, and conduct thorough post-incident reviews.
- Develop advanced automation scripts and tools to enhance system reliability, operational efficiency, and scalability.
- Conduct capacity planning and develop disaster recovery plans to ensure business continuity and future growth.
- Provide technical guidance and mentorship to junior and mid-level engineers to foster their professional growth.
Requirements
- Minimum of 7+ years of experience in site reliability engineering or a closely related field.
- Deep understanding of system reliability concepts, including advanced monitoring, automation, and incident response.
- Proficiency with multiple scripting languages and automation tools.
- Strong experience with cloud platforms and containerization technologies.
- Proven leadership, strategic thinking, and excellent communication skills.
What we offer
- The opportunity to work in a fully remote, autonomous, and agile environment.
- Engagement in strategic planning and decision-making processes for IT infrastructure.
- A culture that values diversity, inclusion, and continuous professional development.
- The chance to influence the direction of IT operations and integrate new technologies.