LLM Pre-training Algorithm Engineer
Responsibilities
- Lead the development of LLM data production pipelines covering sourcing, collection, parsing, processing, experimentation, and analysis, providing stable, large-scale, high-quality pre-training data for foundation models.
- Conduct research on advanced data synthesis, exploring LLM-based data synthesis and augmentation techniques such as Self-Instruct and Agent interaction simulation, and design efficient generation strategies to address data gaps.
- Build automated evaluation systems for synthetic data, such as Reward Models and LLM-as-a-Judge, and use model evaluation and data analysis feedback to iteratively improve production pipelines and data generation strategies.
- Build and optimize the data engineering foundation for LLM pre-training, developing automated frameworks and platforms for large-scale data cleaning, deduplication, and formatting, while improving resource scheduling and data strategy iteration efficiency.
- Curate high-quality pre-training data from across the web, establish end-to-end systems for data quality, diversity, and scenario-specific labeling, and collaborate with algorithm and infrastructure teams to explore the optimal mix of real and synthetic data.
Qualifications
- Master’s degree or above in Computer Science, Artificial Intelligence, Mathematics, or a related field; strong programming fundamentals; proficiency in Python and at least one of Java, Go, or C++.
- Experience in AI data development, understanding of LLM pre-training fundamentals, and familiarity with the data characteristics of at least one core scenario such as code generation or general NLP.
- Professional expertise in either data synthesis or foundational engineering. Research track: deep understanding of mainstream LLM architectures and training mechanisms, familiarity with Prompt techniques and data augmentation methods, and in-depth understanding of data construction for LLM alignment such as RLHF/DPO. Engineering track: familiarity with distributed computing frameworks such as Spark, Flink, and Ray, experience in end-to-end large-scale data cleaning and processing, with familiarity with inference acceleration frameworks such as vLLM preferred.
- Ability to design data quality metrics and use machine learning algorithms to improve data filtering and evaluation efficiency; strong communication skills with the ability to accurately align on requirements and coordinate resources.
Preferred Qualifications
- Experience preparing data for LLMs or successfully training and deploying LLMs using synthetic data.
- High-quality publications in NLP or LLM-related venues such as ACL, EMNLP, or NeurIPS, or active contributions to GitHub open-source projects, particularly in synthetic data or data processing.