Tether Operations Limited is seeking an AI Research Engineer to join our global team and drive innovation in model compression and efficient deployment for advanced multimodal AI systems. You will focus on reducing model footprint and computational costs for large language and vision-language models, enabling high-performance AI to run efficiently on resource-constrained edge devices.
Key responsibilities
- Apply low-bit quantization to reduce model size and inference latency for generative AI models while maintaining high output quality.
- Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models for efficient multimodal reasoning.
- Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead.
- Analyze trade-offs between model efficiency and accuracy, proposing improvements based on empirical findings.
- Author technical papers and publish findings in top-tier conferences to advance the field of model compression.
Requirements
- PhD in NLP, Machine Learning, or a related field, with a strong track record in AI R&D and publications in A* conferences.
- Extensive experience with PyTorch or equivalent deep learning frameworks.
- Hands-on expertise with model quantization, including both Quantization-Aware Training and Post-Training Quantization.
- Proven research and practical experience with knowledge distillation and model pruning.
- Solid understanding of transformer architectures, backpropagation, optimization, and fine-tuning techniques.
What we offer
- The opportunity to work within a global talent powerhouse on cutting-edge AI and blockchain technology.
- A collaborative environment where you can push boundaries and set new industry standards.
- The chance to contribute to innovative platforms that impact millions of users worldwide.