CloudLinux is seeking a Lead Site Reliability Engineer to join our team and take ownership of the reliability platform for Imunify360, our multi-layer Linux server security suite. You will be responsible for defining and implementing a robust SRE framework for approximately 70 components, ensuring that service health is measurable, detectable, and actionable across our entire product line.
Key responsibilities
- Define and implement Service Level Indicators (SLIs) and Objectives (SLOs) for ~70 components in collaboration with squad leads.
- Design and build a scalable telemetry collection pipeline to monitor fleet-wide performance and security control efficacy.
- Develop symptom-based, SLO-anchored alerting systems with clear escalation paths and runbooks.
- Establish incident management practices, including blameless postmortems and on-call rotations for owning squads.
- Ensure that all security controls are effectively enabled and current, making failure detectable in hours rather than months.
Requirements
- Substantial experience in production engineering or SRE, specifically in defining SLO frameworks from scratch.
- Strong proficiency in Python, with the ability to read and modify Go or Rust code.
- Deep practical knowledge of time-series telemetry (Prometheus, Grafana, ClickHouse) and distributed systems debugging.
- Experience with configuration management and CI/CD tools like Ansible and GitLab CI.
- Excellent written communication skills, capable of driving consensus across engineering teams in an asynchronous environment.
What we offer
- Fully remote work environment with flexible hours, allowing you to work from anywhere worldwide.
- Professional development opportunities, including mentorship and knowledge-exchange programs.
- Generous leave policy including 24 days of paid vacation, 10 national holidays, and unlimited sick leave.
- Reimbursement for private medical insurance, co-working spaces, and gym memberships.
- A culture that encourages innovation, including rewards for patentable ideas.