Back

LLM Infrastructure Engineer

Responsibilities

You will work in one or more of the following areas:

  1. Optimize scheduling, computation, and communication for different distributed parallelism strategies in LLM training and inference systems, with the goal of maximizing system performance.
  2. Optimize training or inference kernel performance for different AI accelerator architectures, pushing hardware efficiency and performance to the limit.
  3. Co-design algorithms and systems to achieve the optimal balance between system performance and algorithmic effectiveness.
  4. Build Agent frameworks and platforms to support reinforcement learning model training and performance optimization in complex interactive environments.

Qualifications

  1. Deep understanding of Transformer architectures and core technologies for LLM training and inference, with the ability to analyze their engineering feasibility and bottlenecks.
  2. Strong proficiency in PyTorch, including distributed training and inference, with experience in the design, development, and optimization of large-scale distributed systems.
  3. Good software development practices and solid fundamentals in computer architecture, operating systems, and computer networks.
  4. Strong learning ability and a collaborative team spirit.

Preferred Qualifications

  1. Experience contributing to the development of core frameworks such as Megatron-LM, Transformer Engine, vLLM, or SGLang.
  2. Active participation in the open-source community, with GitHub projects you are proud of.
  3. Familiarity with large-scale AI workload management on cloud platforms.
Deliver