Buzz Solutions is at the forefront of revolutionizing power grid infrastructure maintenance through advanced AI and computer vision technology. We are seeking an Applied Machine Learning Platform Engineer to join our team and enhance the cloud infrastructure, databases, and tooling that power our large-scale computer vision models.
Key responsibilities
- Design, build, and maintain scalable training infrastructure specifically tailored for computer vision workloads.
- Implement and manage distributed training pipelines, including multi-GPU and multi-node setups, to support efficient model training and hyperparameter tuning.
- Develop and maintain robust data pipelines and storage strategies for managing large training datasets, annotations, and model artifacts.
- Implement feature stores, data versioning, and experiment tracking systems to ensure reliable and reproducible model iteration.
- Automate existing analysis workflows and maintain clear documentation for platform components and deployment processes.
Requirements
- 2-4 years of professional experience in platform, backend, data, or MLOps engineering roles.
- Strong proficiency in Python, including idiomatic code, type hints, async patterns, and performance-aware implementation.
- Solid software engineering fundamentals with a focus on testing, code reviews, and component-level system design.
- Hands-on experience operating distributed cloud machine learning infrastructure on platforms like AWS, GCP, or Kubernetes.
- Experience with database design and data systems, including schema design and query optimization for large-scale datasets.
What we offer
- The opportunity to work within a team of experienced ML engineers on high-impact infrastructure projects.
- High level of autonomy to drive your own projects while receiving support for professional growth.
- A fully remote work environment that prioritizes clear communication and technical excellence.