Skip to content

other

GDPO

Multi-reward variant of group-relative policy optimization, described in arXiv:2601.05242, that normalizes each reward channel separately before combining them.

Current clusters