RITS is seeking a skilled Service Manager/Site Reliability Engineer to join a global technology organization. This role focuses on ensuring the reliability, stability, and operational excellence of large-scale production systems by bridging the gap between incident operations, SRE practices, and technical stakeholder communication.
Key responsibilities
- Monitor, triage, and coordinate responses to production incidents and operational alerts.
- Act as a central coordination point between engineering teams and key stakeholders during incidents.
- Manage incident lifecycles from detection through resolution and post-incident activities.
- Support incident reporting, root cause analysis, and operational reviews.
- Contribute to process improvements, automation initiatives, and operational tooling enhancements.
Requirements
- 5+ years of experience in Incident Operations, Site Reliability Engineering, or a similar technical role.
- Proven ability to operate effectively during high-pressure, real-time incident scenarios.
- Experience with monitoring and alerting tools such as Datadog or Chronosphere.
- Hands-on programming experience with Python and/or Kotlin.
- Strong understanding of distributed systems and production environments.
What we offer
- Opportunity to work in a fast-paced, highly available global environment.
- Collaboration with cross-functional engineering and business teams.
- Focus on continuous improvement, automation, and system scalability.
- Support for operational excellence and reliability-focused development.