Svitla Systems is seeking a GPU & ML Infrastructure Engineer to join a technology startup focused on developing physics-informed foundational models for GPU and compute health monitoring. You will play a critical role in building the infrastructure that generates high-quality, reproducible hardware telemetry data used to predict semiconductor degradation and hardware performance.
Key responsibilities
- Port and adapt standardized test procedures to various hardware platforms, including datacenter GPUs and edge devices.
- Maintain and expand a benchmark workload suite, including synthetic kernels and LLM inference/training tasks.
- Ensure logger consistency and data integrity across diverse hardware sources, including in-band and out-of-band telemetry.
- Automate the deployment of workloads and data collection processes to ensure fully reproducible, unattended runs.
- Develop automated data quality checks to identify sampling anomalies and ensure reliable storage handoffs.
Requirements
- Proven experience deploying LLM inference and training workloads on GPUs, including quantized models.
- Strong ability to diagnose sensor and sampling issues within time-series hardware data.
- Deep understanding of Linux systems, GPU driver stacks, process orchestration, and scheduling.
- Expertise in collecting hardware telemetry programmatically via NVML, DCGM, BMC, IPMI, or Redfish.
- Ability to control sources of nondeterminism to ensure workload repeatability.
What we offer
- Opportunity to work on cutting-edge physics-informed AI models for hardware health.
- Full ownership of the end-to-end data generation pipeline for a growing fleet of GPU targets.
- Collaborative environment focused on high-precision systems research and infrastructure engineering.