80 papers from 2025–2026 on what GRPO optimizes, where its signal fails, how researchers preserve diversity and reduce rollout cost, and when simpler methods win · 80 papers · 中文 ↗
Selection rule. This re-audit keeps papers whose central contribution explains or changes GRPO, identifies a missing failure mode, or gives a domain-specific counterexample. Papers that merely use stock GRPO are excluded. The resulting structure follows the research sequence: establish the objective and failure, fix gradient or credit, validate rewards, preserve exploration, allocate and reuse rollouts, then test whether the same assumptions survive in agents, generation, and recommendation.
0 / 80 read
How to judge a new variant. First hold the base model, prompt template, reward, and verifier fixed. Then name the exact failure: objective bias, silent groups, poor token credit, noisy reward, entropy collapse, wasted rollouts, policy lag, or a domain-specific mismatch. Change only the component tied to that failure and compare generated tokens, verifier calls, update compute, wall-clock time, Pass@1, Pass@K, and out-of-domain quality. If a filtered REINFORCE, global normalization, reward-weighted SFT, or critic-based PPO baseline wins under the same budget, use the simpler method.