LLM Training Platform Engineer (MLOps / AI Infra) — The Hong Kong Polytechnic University
Responsibilities
- Design and develop LLM training platforms, building capabilities for GPU resource management, training job scheduling, inference services, and MLOps to support efficient LLM training and iteration.
- Build and optimize GPU clusters using Kubernetes + NVIDIA GPU Operator, including resource scheduling, container environments, and configuration and optimization of the GPU software stack such as CUDA, Driver, and NCCL.
- Build core training platform toolchains, including training job orchestration, automated pipelines, base image management, data loading, and model version management.
- Collaborate with algorithm teams to support and optimize distributed training frameworks such as PyTorch Distributed, Megatron, and SGLang, improving training efficiency and stability.
- Build platform monitoring and automated operations systems, improving monitoring, alerting, and troubleshooting capabilities across GPUs, networks, storage, and training jobs.
Qualifications
- Master’s degree or above in Computer Science, Communications, Electronics, or a related field.
- More than 5 years of backend development experience, or more than 3 years of experience in distributed systems or large-scale computing platform development.
- Familiarity with LLM training workflows and understanding of the complete training, inference, and evaluation lifecycle.
- Proficiency in programming languages such as Python / Go, with strong engineering development skills.
- Familiarity with Kubernetes, Docker, GPU scheduling, and AI computing environment construction.
- Familiarity with NVIDIA GPU architecture, CUDA, NCCL, InfiniBand, and other high-performance computing technologies.
Preferred Qualifications
- Experience building GPU training clusters at the thousand-GPU scale.
- Familiarity with LLM training and inference frameworks such as Ray, DeepSpeed, PyTorch, Megatron, vLLM, and SGLang.
- Experience building GPU training platforms, MLOps platforms, or distributed systems.
- Familiarity with large-scale storage systems such as HDFS, GPFS, and JuiceFS.