A Two-Year-Old Algorithm Finally Meets Its Limits
GRPO has been the default critic-free RL recipe since DeepSeekMath introduced it in early 2024 [7], and it went on to train DeepSeek-R1 [8]. Its core trick - replacing a learned value function with a group-relative baseline - works well for single-turn reasoning tasks like math problems, where one final answer gets one reward. But as researchers pushed it into long-horizon agent settings - clicking through a web shop, navigating a simulated house, chaining tool calls across dozens of steps - the same trick breaks down. ProVer states the problem plainly: GRPO's 'uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success' [3].
What's notable is the timeline. PhGPO [5]and SPADER [6]flagged this exact gap months earlier in 2026, applying fixes to tool-planning and multi-answer QA respectively. Then, in a much tighter four-day window - September 28 to October 1, 2026 - four more independent teams converged on the same diagnosis for the same two embodied/web benchmarks: ProVer on September 28 [3], SHARPO and T2SPO both on September 30 [1][2], and DARS on October 1 [4]. That clustering makes this look less like a single breakthrough and more like a sign that the field quietly hit the same wall at roughly the same time.


