Back

LLM Post-training Algorithm Engineer

Responsibilities

  1. End-to-end alignment optimization: develop the full post-training pipeline from SFT to RL to improve model performance in general dialogue, complex reasoning, and Agentic tasks.
  2. Data system development: drive data requirements through model capability diagnosis and design high-quality post-training data pipelines covering the cleaning, synthesis, and evaluation of reasoning data and Agent trajectory data.
  3. Advanced capability development: optimize data and algorithms for complex reasoning and long-horizon Agent tasks, including multi-step planning, tool use, and environment interaction, to push model capability boundaries.

Qualifications

  1. Deep understanding of Transformer architectures and the full LLM training lifecycle; strong understanding of post-training algorithms such as SFT, PPO, and GRPO; familiarity with Agentic RL paradigms such as RLVR and multi-step trajectory rewards.
  2. Strong data insight, with the ability to define high-quality data standards; proficiency in large-scale data cleaning, deduplication, diversity control, and quality evaluation; familiarity with constructing Chain-of-Thought (CoT) and tool-use trajectory data.
  3. Proficiency in PyTorch, hands-on experience with large-scale distributed training, and familiarity with RL training frameworks such as veRL, slime, and OpenRLHF.
  4. Strong bad-case analysis skills and the ability to identify model bottlenecks through fine-grained decomposition; strong learning ability, clear logic, effective collaboration, and the ability to track and translate cutting-edge research into practice.

Preferred Qualifications

  1. End-to-end LLM R&D experience and experience releasing or deploying high-quality models.
  2. Familiarity with code, mathematics, Agent, and other reasoning tasks; understanding of automated verification such as Sandbox / Formal Verification and execution-feedback mechanisms; familiarity with Agent benchmarks such as SWE-bench and τ-bench.
  3. Publications at top-tier conferences such as NeurIPS, ICLR, ICML, or ACL, or high-level competition awards.
Deliver