Skip to content

other

Group-Sequence Policy Optimization

Reinforcement learning method that scores entire action sequences at the end of a multi-step task rather than rewarding individual token predictions.

Known aliases

  • GSPO

Relationships

No evidence-backed relationships are recorded.

Current clusters