RITS is seeking a skilled Python Developer to join our team in a Site Reliability Engineering (SRE) capacity. You will play a critical role in ensuring the reliability, stability, and operational excellence of large-scale production systems within a global technology organization.
Key responsibilities
- Monitor, triage, and coordinate responses to production incidents and operational alerts.
- Act as a central coordination point between engineering teams and stakeholders during incidents.
- Manage incident lifecycles from detection through resolution and post-incident activities.
- Contribute to process improvements, automation initiatives, and operational tooling enhancements.
- Collaborate with engineering teams to improve observability, monitoring, and incident response capabilities.
Requirements
- 5+ years of experience in Software Engineering, SRE, or Production Engineering roles.
- Strong background in building and maintaining production-grade applications using Python.
- Proven ability to operate effectively during high-severity, real-time production incidents.
- Solid understanding of distributed systems, cloud-native architectures, and large-scale environments.
- Experience with software engineering best practices, including testing, CI/CD, and version control.
What we offer
- Fully remote work environment within a global technology organization.
- Opportunity to drive automation and operational excellence in high-availability systems.
- Engagement in complex technical investigations and root cause analysis.