itMatch is looking for a Senior Site Reliability Engineer to join a team focused on solving complex reliability challenges for a cutting-edge serverless inference platform. You will play a key role in building automation, scaling systems, and contributing to the architecture of AI-driven infrastructure.
Key responsibilities
- Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, and SLO/SLI tracking.
- Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
- Integrating AI workloads into incident management processes, building runbooks, and conducting blameless post-mortems.
- Collaborating with product engineering teams to improve reliability and ensure operational readiness for product releases.
- Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.
Requirements
- Extensive operational experience in SRE, infrastructure, or platform engineering managing large-scale distributed systems.
- Deep expertise in Kubernetes and large-scale containerization systems.
- Proficiency in Python or Go for automation, CI/CD pipelines, and infrastructure-as-code tools like Terraform.
- Experience defining SLOs and utilizing observability tools such as Prometheus and Grafana.
- Strong interest in or experience with AI/ML infrastructure, model serving, or GPU-based workloads.
What we offer
- Opportunities to grow, flourish, and achieve professional goals within a supportive environment.
- Comprehensive benefit options designed to support your health, finances, family, and personal time.
- A flexible approach to benefits that can be tailored to meet your individual needs and budget.