Link Group is seeking a highly skilled Senior Site Reliability Engineer to join a critical team responsible for a massive global data pipeline. You will play a pivotal role in ensuring the reliability, performance, and integrity of large-scale distributed data systems that process telemetry from hundreds of thousands of servers.
Key responsibilities
- Engineer and maintain the reliability of massive data platforms, focusing on distributed databases like Cassandra and MongoDB.
- Act as the primary technical escalation point for complex reliability and performance issues across the global infrastructure.
- Design and implement advanced observability solutions using Prometheus and Grafana to monitor platform health.
- Develop high-quality automation and internal tooling using Python, Go, or Java to streamline operations and diagnostics.
- Collaborate with Engineering and Product teams to influence architecture and deliver scalable solutions that prevent systemic failures.
Requirements
- At least 5 years of experience as an SRE or Systems Engineer managing mission-critical distributed systems.
- Deep hands-on experience with large-scale NoSQL databases, specifically Cassandra or MongoDB.
- Proficiency in programming languages such as Python, Go, or Java for building automation and tooling.
- Strong foundation in Linux/Unix administration and low-level system troubleshooting.
- Proven ability to identify root causes of complex technical problems and implement robust, production-grade fixes.
What we offer
- Opportunity to work on global-scale infrastructure with a focus on high-performance data systems.
- A collaborative environment where you can influence architectural decisions and long-term reliability strategies.
- Access to modern observability and automation tools to solve complex engineering challenges.