Back

LLM Pre-training Algorithm Engineer

Responsibilities

  1. Lead the development of LLM data production pipelines covering sourcing, collection, parsing, processing, experimentation, and analysis, providing stable, large-scale, high-quality pre-training data for foundation models.
  2. Conduct research on advanced data synthesis, exploring LLM-based data synthesis and augmentation techniques such as Self-Instruct and Agent interaction simulation, and design efficient generation strategies to address data gaps.
  3. Build automated evaluation systems for synthetic data, such as Reward Models and LLM-as-a-Judge, and use model evaluation and data analysis feedback to iteratively improve production pipelines and data generation strategies.
  4. Build and optimize the data engineering foundation for LLM pre-training, developing automated frameworks and platforms for large-scale data cleaning, deduplication, and formatting, while improving resource scheduling and data strategy iteration efficiency.
  5. Curate high-quality pre-training data from across the web, establish end-to-end systems for data quality, diversity, and scenario-specific labeling, and collaborate with algorithm and infrastructure teams to explore the optimal mix of real and synthetic data.

Qualifications

  1. Master’s degree or above in Computer Science, Artificial Intelligence, Mathematics, or a related field; strong programming fundamentals; proficiency in Python and at least one of Java, Go, or C++.
  2. Experience in AI data development, understanding of LLM pre-training fundamentals, and familiarity with the data characteristics of at least one core scenario such as code generation or general NLP.
  3. Professional expertise in either data synthesis or foundational engineering. Research track: deep understanding of mainstream LLM architectures and training mechanisms, familiarity with Prompt techniques and data augmentation methods, and in-depth understanding of data construction for LLM alignment such as RLHF/DPO. Engineering track: familiarity with distributed computing frameworks such as Spark, Flink, and Ray, experience in end-to-end large-scale data cleaning and processing, with familiarity with inference acceleration frameworks such as vLLM preferred.
  4. Ability to design data quality metrics and use machine learning algorithms to improve data filtering and evaluation efficiency; strong communication skills with the ability to accurately align on requirements and coordinate resources.

Preferred Qualifications

  1. Experience preparing data for LLMs or successfully training and deploying LLMs using synthetic data.
  2. High-quality publications in NLP or LLM-related venues such as ACL, EMNLP, or NeurIPS, or active contributions to GitHub open-source projects, particularly in synthetic data or data processing.
Deliver