Research Plan

GRPO after the Boom

80 papers from 2025–2026 on what GRPO optimizes, where its signal fails, how researchers preserve diversity and reduce rollout cost, and when simpler methods win · 80 papers · 中文 ↗
Selection rule. This re-audit keeps papers whose central contribution explains or changes GRPO, identifies a missing failure mode, or gives a domain-specific counterexample. Papers that merely use stock GRPO are excluded. The resulting structure follows the research sequence: establish the objective and failure, fix gradient or credit, validate rewards, preserve exploration, allocate and reuse rollouts, then test whether the same assumptions survive in agents, generation, and recommendation.
0 / 80 read
How to judge a new variant. First hold the base model, prompt template, reward, and verifier fixed. Then name the exact failure: objective bias, silent groups, poor token credit, noisy reward, entropy collapse, wasted rollouts, policy lag, or a domain-specific mismatch. Change only the component tied to that failure and compare generated tokens, verifier calls, update compute, wall-clock time, Pass@1, Pass@K, and out-of-domain quality. If a filtered REINFORCE, global normalization, reward-weighted SFT, or critic-based PPO baseline wins under the same budget, use the simpler method.