The Wikimedia Foundation is seeking a Senior Site Reliability Engineer to join our team and help maintain the infrastructure powering Wikipedia’s data feeds. You will play a critical role in designing, developing, and scaling reliable API services that support the global knowledge ecosystem.
Key responsibilities
- Define, track, and improve Service Level Objectives (SLOs) and error budgets to ensure high availability.
- Build and enhance observability systems, including metrics, logs, and distributed tracing.
- Design and optimize CI/CD and GitOps workflows to enable automated, reliable deployments.
- Partner with engineering teams to embed reliability best practices throughout the development lifecycle.
- Participate in incident response and on-call rotations to maintain system health.
Requirements
- Proven experience operating highly available, large-scale distributed systems.
- Proficiency in at least one programming language such as Python or Go.
- Strong experience with Infrastructure as Code tools like Terraform or Ansible.
- Deep understanding of cloud infrastructure (AWS, Azure, or GCP) and scalability principles.
- Excellent communication skills and the ability to work effectively in a distributed, cross-functional team.
What we offer
- The opportunity to work on high-impact projects that serve billions of users worldwide.
- A collaborative, remote-first environment with a diverse and global team of engineers.
- A culture that values continuous improvement, blameless postmortems, and technical innovation.
- The chance to contribute to a mission-driven organization dedicated to free knowledge.