The Wikimedia Foundation is seeking a Senior Site Reliability Engineer to help maintain and evolve the infrastructure powering Wikipedia, one of the world's most visited websites. As part of our globally distributed SRE team, you will work in an open-source environment to ensure the reliability, scalability, and security of our global platform.
Key responsibilities
- Manage day-to-day operations, deployment, and configuration of public-facing infrastructure.
- Utilize configuration management tools like Puppet and Kubernetes to automate service maintenance.
- Collaborate with product teams on architectural design to support scalable functionality.
- Participate in a 24/7 on-call rotation, including incident response and root cause analysis.
- Mentor peers and contribute to a culture of continuous improvement and asynchronous collaboration.
Requirements
- 6+ years of experience in SRE, Operations, or DevOps roles.
- Proficiency in Python and shell scripting, with strong Linux system-level troubleshooting skills.
- Experience managing infrastructure security for large, diverse service fleets.
- Strong communication skills in English and the ability to work effectively across time zones.
- Proven history of automating processes and identifying technical gaps.
What we offer
- The opportunity to work on one of the most impactful open-source projects in the world.
- A remote-first work environment with a diverse, global team.
- Engagement with a mission-driven organization dedicated to free knowledge.