Gradle is looking for founding members to join our new Site Reliability Engineering team. You will play a critical role in ensuring the reliability, performance, and availability of Develocity, an observability and intelligence platform used by leading global software organizations.
Key responsibilities
- Operate and maintain all Develocity instances and supporting services.
- Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting across the stack.
- Drive automation across application deployment, upgrades, monitoring, and recovery.
- Build and maintain observability for all managed services using logging, metrics, and tracing.
- Collaborate with engineering teams to integrate reliability into features from the start.
Requirements
- 5+ years of experience in SRE, DevOps, or an equivalent role operating production services at scale.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, specifically with AWS services like EKS, RDS, and S3.
- Proficiency with observability tools such as Prometheus and Grafana, and Infrastructure as Code using Terraform.
- Strong written and verbal English communication skills for a distributed, remote-first environment.
What we offer
- A ground-floor opportunity to shape SRE practices within a new team.
- Real ownership of production systems used by world-class software engineering organizations.
- A culture that prioritizes automation and thoughtful engineering over manual heroics.
- The ability to work fully remotely from anywhere in Europe within the GMT timezone.