Chess.com is the world's leading platform for chess, serving a global community of over 250 million players. We are a mission-driven, fully remote team of over 600 people dedicated to providing the best tools and content for chess enthusiasts worldwide.
Key responsibilities
- Design and implement multi-regional resilient infrastructure capable of handling millions of concurrent sessions daily.
- Lead hybrid cloud migration strategies, integrating bare-metal datacenter resources with cloud services for optimal performance.
- Own on-call rotation and incident response procedures to maintain high availability and rapid resolution of critical issues.
- Architect monitoring and alerting systems to proactively identify and resolve performance bottlenecks.
- Collaborate with development teams to implement infrastructure-as-code practices and robust deployment pipelines.
Requirements
- 5+ years of experience in site reliability engineering, DevOps, or infrastructure engineering roles.
- Strong proficiency with UNIX/Linux operating systems and command-line administration.
- Experience with cloud platforms like GCP, AWS, or Azure and infrastructure-as-code tools such as Terraform.
- Hands-on experience with configuration management systems and containerization technologies like Docker or Kubernetes.
- Solid understanding of networking fundamentals, protocols, and network troubleshooting.
- Demonstrated sense of ownership and accountability for system reliability and performance.
What we offer
- A full-time position within a global, mission-driven organization.
- The opportunity to work 100% remotely from anywhere in the world.
- A collaborative, flat, and life-celebrating culture that values innovation and technical excellence.