Valtech is an experience innovation company and a trusted partner to the world’s most recognized brands. As a Senior Site Reliability Engineer, you will act as a bridge between software development and operations, helping our teams deliver reliable, high-speed digital solutions while maintaining an excellent customer experience.
Key responsibilities
- Collaborate with teams to define and implement SLIs and SLOs.
- Design and maintain robust systems for observability.
- Analyze failure scenarios and develop effective mitigation strategies.
- Create and maintain runbooks to prevent or remediate production issues.
- Facilitate incident management processes, including participation in on-call rotations.
Requirements
- At least 5 years of experience in software, DevOps, or cloud engineering, with at least 2 years specifically as an SRE.
- Strong experience with incident management in high-traffic, 24x7 production environments.
- Proficiency in cloud platforms (AWS, Azure, or GCP) and serverless services.
- Extensive knowledge of monitoring tools like Datadog, New Relic, or Prometheus/Grafana.
- Experience with CI/CD pipelines (GitHub, Azure DevOps, Jenkins) and microservices (Docker, Kubernetes).
- Excellent communication skills in English (C1 level or above).
What we offer
- 24 working days of paid vacation plus national holidays.
- Comprehensive medical insurance and wellness benefits like Multisport.
- Access to internal workshops, professional certifications, and a structured mentoring program.
- A supportive, inclusive global culture that values autonomy and continuous learning.
- A progressive benefits package that grows with your tenure at the company.