Skip to content

other

Group relative policy optimization

Reinforcement learning method that scores a group of sampled model responses against each other to update the policy without a separate value network.

Current clusters