build1 distinct publisher
Outcome-only GRPO squeezed a toy model's answers onto about 2.5 of 99 possible sums
Greedy accuracy came back to baseline by 1,500 steps and per-sample correctness tripled, so the fall in pass@64 to 0.19 only surfaced when the authors paid for 64 samples a problem instead of one.
Publishers:dev.to
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+18
- Incentives55