RemoDevs is seeking an experienced Senior Site Reliability Engineer or DevOps professional to join our team. You will play a critical role in managing highly available, production-grade cloud and Kubernetes environments while advocating for SRE best practices and building secure, scalable CI/CD pipelines.
Key responsibilities
- Promote core SRE principles including monitoring, alerting, logging, tracing, SLOs, and toil reduction.
- Administer high-availability Kubernetes clusters and cloud infrastructure across multiple environments.
- Provision and manage cloud and on-premise infrastructure using Terraform and Terragrunt.
- Build, secure, and optimize automated CI/CD workflows using tools like GitHub Actions or Concourse.
- Manage data-related platform technologies such as Kafka, Airflow, and Hadoop while supporting platform engineering.
Requirements
- Proven professional experience working as an SRE or DevOps Engineer.
- Strong hands-on experience with Kubernetes, ArgoCD, Istio, and Helm.
- Proficiency in Infrastructure as Code tools like Terraform, Terragrunt, Ansible, or Salt.
- Demonstrated expertise in setting up and optimizing CI/CD pipelines and observability stacks like Prometheus, Grafana, or NewRelic.
- Excellent written and spoken English skills with a strong ability to collaborate and mentor software development teams.
What we offer
- Opportunity to work in a fully remote environment with a focus on modern cloud-native technologies.
- Engagement in challenging projects involving high-scale Kubernetes clusters and complex data infrastructure.
- Collaborative culture that values SRE best practices and continuous improvement of automated workflows.