emagine Polska is seeking a Senior DevOps / Site Reliability Engineer to ensure the reliability, scalability, and security of our cloud-native platform. You will play a pivotal role in optimizing our infrastructure, implementing SRE best practices, and supporting development teams in delivering highly available services.
Key responsibilities
- Design, implement, and maintain highly available and scalable infrastructure on AWS.
- Own and improve the reliability of production systems using SRE principles such as SLOs, SLIs, and error budgets.
- Manage and optimize container orchestration platforms including Kubernetes, Docker, and Helm.
- Lead incident response, perform root cause analysis, and drive continuous improvement through postmortems.
- Implement and maintain comprehensive monitoring, logging, and alerting solutions.
Requirements
- Over 5 years of professional experience in DevOps, SRE, or Platform Engineering.
- Proven expertise in managing Kubernetes in production environments.
- Strong proficiency in Infrastructure as Code, with a preference for Terraform.
- Advanced proficiency in the French language.
- Solid experience with CI/CD pipelines and Linux systems administration.
What we offer
- Opportunity to work in a fully remote environment on complex, large-scale cloud projects.
- Collaborative culture focused on technical excellence and SRE best practices.
- Professional development within a team of experienced engineers.