ProVer fixes GRPO's credit assignment by verifying only pivotal decisions
Fixing GRPO's credit assignment problem without evaluating every step
GRPO assigns the same trajectory-level advantage to every token, blurring which decisions actually mattered. ProVer uses an agentic judge to contrast successful and failed rollouts and propose a segment that may explain the difference. Instead of trusting the judge, it verifies the segment by comparing terminal success rates of continuations sampled before and after it, then folds positive estimates into GRPO advantages. On ALFWorld, WebShop, and SearchQA, ProVer beats GRPO by 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B.
By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state.