Tether is a global leader in digital finance, pioneering innovative solutions across blockchain technology, energy, and AI. We are looking for an AI Research Engineer to join our team and drive advancements in model serving and inference architectures for high-performance, scalable AI systems.
Key responsibilities
- Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage.
- Build, run, and monitor controlled inference tests in simulated and live production environments to track performance metrics.
- Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges on resource-constrained devices.
- Analyze computational efficiency and diagnose bottlenecks in the serving pipeline to ensure scalability and reliability.
- Collaborate with cross-functional teams to integrate optimized inference frameworks into production pipelines for edge and on-device applications.
Requirements
- Degree in Computer Science or a related field, with a strong track record in AI R&D.
- Expertise in Metal Shading Language (MSL) and writing custom compute shaders from scratch.
- Proven experience in low-level kernel optimizations and inference optimization on mobile devices.
- Deep understanding of modern model serving architectures, inference techniques, and GPU kernel development.
- Knowledge of distributed inference systems, Diffusion Models, Vision Transformers, and optimization methods like quantization and pruning.
What we offer
- The opportunity to work within a global talent powerhouse on cutting-edge AI and blockchain projects.
- A fully remote work environment allowing collaboration with brilliant minds from across the world.
- The chance to push the boundaries of AI performance and contribute to industry-leading financial and data technologies.