LLM Evaluation Algorithm Engineer
Responsibilities
- Participate in the continuous development and iteration of the LLM evaluation framework, covering evaluation execution, Benchmark integration, result aggregation, automated uploads, and visualization analysis, improving system stability, usability, and scalability.
- Build evaluation capabilities for Agent, Code, mathematics, knowledge, reasoning, PPL, and other scenarios, supporting rule-based evaluation, LLM-as-a-Judge, PPL/logprobs evaluation, multi-turn dialogue evaluation, and other paradigms.
- Continuously optimize evaluation task scheduling and automation workflows; participate in evaluation visualization platform development and improve ETL, DuckDB/Parquet storage, frontend/backend APIs, sample-level analysis, cross-model comparison, and trend tracking.
- Track mainstream open-source and industry Benchmarks; clean, adapt, version, and validate evaluation datasets and metrics to ensure fairness, reproducibility, and interpretability.
- Analyze model evaluation results in depth, identify strengths and weaknesses across knowledge, reasoning, code, mathematics, instruction following, long context, and Agent capabilities, and provide feedback for training iteration and data improvement.
- Collaborate with training, inference, data, and product teams to build a closed loop from model output to automated evaluation, result feedback, issue diagnosis, and capability improvement.
Qualifications
- Bachelor’s degree or above; Computer Science, Artificial Intelligence, Software Engineering, Mathematics, or related majors preferred.
- Proficiency with common development and cluster tools including Linux, Shell, Git, Docker/Singularity, and Slurm.
- Familiarity with common LLM Benchmarks and evaluation methods such as MMLU, CMMLU, CEval, BBH, GPQA, GSM8K, HumanEval, LiveCodeBench, SuperGPQA, and PPL.
- Strong system design capabilities and the ability to design reliable solutions for evaluation task scheduling, concurrent execution, result storage, log tracing, failure retries, and visualization analysis.
- Strong problem-solving skills and the ability to identify evaluation anomalies, performance bottlenecks, and result deviations from logs, metrics, samples, configurations, and code paths.
- Strong interest in LLM evaluation, model capability analysis, and exploring the boundaries of Agent/Code/reasoning capabilities, with willingness to continuously follow industry developments and open-source Benchmarks.
Preferred Qualifications
- Experience using or extending evaluation frameworks such as EvalScope, OpenCompass, lm-evaluation-harness, HELM, or LightEval.
- Experience with FastAPI, Next.js, DuckDB, Parquet, data platforms, or visualization systems.
- Experience with LLM training, post-training, data construction, model analysis, or online inference services.
- Experience with Code Agents, Agent Benchmarks, tool-use evaluation, or multi-turn task evaluation.
- Open-source contributions, technical blogs, evaluation reports, or Benchmark construction experience.