Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations to maintain and improve the reliability, scalability, and performance of our Java-based applications.
Key responsibilities
- Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment.
- Deploy and manage the Grafana stack to deliver real-time monitoring, logging, and alerting.
- Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage and incident response.
- Develop tool-calling LLM agents that interact with infrastructure APIs to execute diagnostic and remediation actions.
- Support incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence.
Requirements
- 5+ years in SRE, DevOps, or similar infrastructure roles managing large-scale production systems.
- 3+ years hands-on experience managing production Kubernetes clusters.
- Advanced expertise with the Grafana observability stack and PromQL.
- 1+ years of practical experience building or operating AI/LLM-powered tools or agents in a production context.
- Strong scripting abilities in Python, Bash, or Go.
- Familiarity with agentic orchestration frameworks like LangChain, LangGraph, or CrewAI.
What we offer
- The autonomy to experiment with cutting-edge AI tools and deploy them in production.
- A mandate to measurably reduce operational toil through intelligent systems.
- Collaboration with a team that treats AI and agentic automation as core competencies.