Skip to content

other

Group Relative Policy Optimization

GRPO is a reinforcement learning algorithm that fine-tunes language models by normalizing rewards across sampled rollouts, avoiding a separate value network.

Known aliases

  • Group Relative Policy Optimization
  • GRPO
  • GRPO-style RL
  • Soft-GRPO

Relationships

No evidence-backed relationships are recorded.

Current stories

invest4 publishers

PewDiePie says OpenAI banned him twice as he trained his Ajax model on GPT-5.6 Sol outputs

PewDiePie says OpenAI banned him twice while he used GPT-5.6 Sol outputs to train Ajax, a 9-billion-parameter model for home PCs. OpenAI has not commented publicly, but his account shows that a team training its own model on a frontier lab's answers can lose API access partway through the build.

Perspective Coverage

4 publishers
Builder
Builder 45%
Operator
Operator 35%
Investor
Investor 20%

Reality

Evidence45
Adoption5
Hype gap+35
Incentives60
Confidence55
build1 publisher

Crutches built from measured failures lift a local Qwen 3B from 33% to 52% on post-cutoff facts

Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence40