Skip to content

Topic

Reinforcement learning for LLM post-training

Methods such as PPO and GRPO that fine-tune a trained language model by optimizing it against reward signals rather than fixed labels.

Current clusters