Back

LLM Training Platform Engineer (MLOps / AI Infra) — The Hong Kong Polytechnic University

Responsibilities

  1. Design and develop LLM training platforms, building capabilities for GPU resource management, training job scheduling, inference services, and MLOps to support efficient LLM training and iteration.
  2. Build and optimize GPU clusters using Kubernetes + NVIDIA GPU Operator, including resource scheduling, container environments, and configuration and optimization of the GPU software stack such as CUDA, Driver, and NCCL.
  3. Build core training platform toolchains, including training job orchestration, automated pipelines, base image management, data loading, and model version management.
  4. Collaborate with algorithm teams to support and optimize distributed training frameworks such as PyTorch Distributed, Megatron, and SGLang, improving training efficiency and stability.
  5. Build platform monitoring and automated operations systems, improving monitoring, alerting, and troubleshooting capabilities across GPUs, networks, storage, and training jobs.

Qualifications

  1. Master’s degree or above in Computer Science, Communications, Electronics, or a related field.
  2. More than 5 years of backend development experience, or more than 3 years of experience in distributed systems or large-scale computing platform development.
  3. Familiarity with LLM training workflows and understanding of the complete training, inference, and evaluation lifecycle.
  4. Proficiency in programming languages such as Python / Go, with strong engineering development skills.
  5. Familiarity with Kubernetes, Docker, GPU scheduling, and AI computing environment construction.
  6. Familiarity with NVIDIA GPU architecture, CUDA, NCCL, InfiniBand, and other high-performance computing technologies.

Preferred Qualifications

  1. Experience building GPU training clusters at the thousand-GPU scale.
  2. Familiarity with LLM training and inference frameworks such as Ray, DeepSpeed, PyTorch, Megatron, vLLM, and SGLang.
  3. Experience building GPU training platforms, MLOps platforms, or distributed systems.
  4. Familiarity with large-scale storage systems such as HDFS, GPFS, and JuiceFS.
Deliver