Margo is seeking a skilled Network Reliability Engineer to join our team and help build and maintain large-scale AI infrastructure. You will play a critical role in ensuring the stability, scalability, and performance of our GPU and HPC clusters while collaborating with engineering teams to drive continuous improvement.
Key responsibilities
- Build and manage large AI infrastructure, focusing on monitoring, diagnosis, and remediation of production incidents.
- Troubleshoot high-impact production issues and participate in an on-call rotation to ensure service continuity.
- Implement and maintain observability solutions to monitor the health of AI infrastructure and applications.
- Promote best practices regarding system stability, resiliency, scalability, and security.
- Collaborate with development teams to ensure infrastructure readiness and contribute to system evolution.
Requirements
- Strong proficiency in Go or Python, complemented by advanced scripting skills in Bash.
- Hands-on experience with Linux systems (Ubuntu/Debian) and networking protocols such as TCP/IP, BGP, DNS, and load-balancing.
- Experience with Infrastructure-as-Code tools like Ansible or Salt and familiarity with monitoring stacks like Prometheus and Grafana.
- Understanding of CI/CD pipelines and experience managing relational databases like MariaDB.
- Proactive mindset with a passion for automation and strong communication skills in English.
What we offer
- Opportunity to work on cutting-edge AI and GPU infrastructure projects.
- A collaborative environment that values knowledge sharing and professional growth.
- Full autonomy in a remote-first setting with a focus on continuous improvement and technical excellence.