LLM Infrastructure Engineer
Responsibilities
You will work in one or more of the following areas:
- Optimize scheduling, computation, and communication for different distributed parallelism strategies in LLM training and inference systems, with the goal of maximizing system performance.
- Optimize training or inference kernel performance for different AI accelerator architectures, pushing hardware efficiency and performance to the limit.
- Co-design algorithms and systems to achieve the optimal balance between system performance and algorithmic effectiveness.
- Build Agent frameworks and platforms to support reinforcement learning model training and performance optimization in complex interactive environments.
Qualifications
- Deep understanding of Transformer architectures and core technologies for LLM training and inference, with the ability to analyze their engineering feasibility and bottlenecks.
- Strong proficiency in PyTorch, including distributed training and inference, with experience in the design, development, and optimization of large-scale distributed systems.
- Good software development practices and solid fundamentals in computer architecture, operating systems, and computer networks.
- Strong learning ability and a collaborative team spirit.
Preferred Qualifications
- Experience contributing to the development of core frameworks such as Megatron-LM, Transformer Engine, vLLM, or SGLang.
- Active participation in the open-source community, with GitHub projects you are proud of.
- Familiarity with large-scale AI workload management on cloud platforms.