T-Mobile is seeking an experienced AIOps Engineer to join the T-Hub team, focusing on the development and maintenance of robust AI infrastructure. In this role, you will be responsible for orchestrating large-scale inference services and ensuring the reliability and performance of our AI platforms in a production environment.
Key responsibilities
- Design, deploy, and maintain vLLM inference services on OpenShift and Kubernetes running on bare-metal GPU infrastructure.
- Manage NVIDIA GPU resources, including partitioning and allocation, to maximize utilization across multiple models and tenants.
- Automate model lifecycle management, including onboarding, versioning, deployment, and rollback processes.
- Build and maintain comprehensive observability solutions, including metrics collection, logging, and tracing for AI inference services.
- Collaborate with AI Engineering and Platform teams to ensure secure, scalable, and reliable AI service delivery.
Requirements
- 5+ years of experience in DevOps, SRE, or Platform Engineering, with at least 2 years specifically in MLOps or AI infrastructure.
- Strong expertise in Kubernetes and OpenShift administration within production environments.
- Proven experience deploying vLLM-based inference platforms and deep knowledge of GPU technologies and CUDA.
- Strong Python programming skills and proficiency in Linux systems administration and Bash scripting.
- Familiarity with CI/CD tools such as GitLab CI, Jenkins, and ArgoCD, alongside Infrastructure as Code practices.
What we offer
- The opportunity to work at the forefront of modern technologies like 5G, IoT, and AI within a leading telecommunications company.
- Access to professional training, courses, and a supportive environment that encourages innovation and skill development.
- Comprehensive benefits package including private medical care, life insurance, and various social perks.
- Flexible working arrangements and a culture that values diversity, agility, and continuous knowledge exchange.