Arista Networks is seeking a skilled Site Reliability Engineer to join our Engineering Productivity team. You will play a vital role in designing, building, and maintaining the secure, scalable, and fault-tolerant infrastructure that powers our product development teams in a fully remote environment.
Key responsibilities
- Build, deploy, and operate critical production systems with a focus on scalability, reliability, observability, and security.
- Develop automation tools to eliminate manual toil and improve the efficiency of production operations.
- Proactively monitor system health, manage incident response runbooks, and conduct postmortem analyses to prevent recurring issues.
- Collaborate with software engineering teams to identify and resolve infrastructural bottlenecks in their workflows.
- Adopt and implement industry best practices to ensure robust and fault-tolerant platform performance.
Requirements
- Bachelor’s or Master’s degree in Computer Science or Engineering, or equivalent professional experience.
- Proficiency in Go, Python, or shell scripting for implementing automation workflows.
- Strong knowledge of Linux/UNIX administration, debugging, and server provisioning.
- Hands-on experience operating complex software systems and infrastructure at scale.
- Proven problem-solving skills and experience with infrastructure-as-code practices.
What we offer
- Opportunity to work with a diverse, innovative team at an industry leader in data-driven networking.
- Exposure to a wide range of modern technologies including Kubernetes, Grafana, Ansible, and cloud-based systems.
- A culture that values diversity, inclusion, and continuous professional growth.
- A fully remote work environment that prioritizes collaboration and technical excellence.