ByteDance and Tsinghua's DAPO sets new bar for open-source LLM reinforcement learning

DAPO: An Open-Source RL System from ByteDance Seed and Tsinghua Air

ByteDance and Tsinghua's DAPO sets new bar for open-source LLM reinforcement learning

DAPO, a fully open-source RL system from ByteDance Seed and Tsinghua AIR, achieves 50 points on AIME 2024 with a Qwen2.5-32B base model, surpassing DeepSeek-R1-Zero-Qwen-32B using only half the training steps. The release includes the algorithm, code infrastructure, dataset, and model weights, providing practical access to scalable reinforcement learning for the research community.

DAPO achieves 50 points on AIME 2024 based on the Qwen2.5-32B base model, outperforming the previous SoTA DeepSeek-R1-Zero-Qwen-32B with 50% training steps.

More from this day

2026-09-21