FYUL is the engine powering on-demand commerce at a global scale, helping creators and brands turn ideas into products. We are looking for a Senior Site Reliability Engineer to join our Platform Infrastructure team and help evolve our container platform and cloud foundation.
Key responsibilities
- Architect and manage highly available, secure, and scalable infrastructure across AWS environments using infrastructure as code.
- Design and operate Amazon EKS clusters, including networking policies, storage, and scaling strategies.
- Drive large-scale automation projects using Terraform and GitOps practices to reduce manual operational work.
- Lead incident response, participate in on-call rotations, and improve system observability using the Grafana stack.
- Mentor mid-level engineers and partner with product squads to ensure reliable and cost-efficient service delivery.
Requirements
- Extensive experience in production infrastructure, specifically with AWS, Kubernetes (EKS), and Linux systems administration.
- Proficiency in Terraform and GitOps workflows, with a strong background in scripting (Python).
- Deep understanding of observability tools like Prometheus, Grafana, and Loki for monitoring and distributed tracing.
- Proven track record of leading complex technical initiatives, incident management, and writing clear documentation.
- Strong communication skills with the ability to mentor team members and collaborate across engineering squads.
What we offer
- A global, inclusive, and supportive team environment.
- Flexible working hours to help you balance your professional and personal life.
- Private health insurance coverage.
- Additional paid days off for mental well-being and personal celebrations.
- Access to continuous learning opportunities, internal meetups, and professional development.