Skip to content

Topic

Reinforcement-Learning Post-Training

Capability gains obtained by extending RL, task environments and post-training compute on an existing base model, using stacks such as IndexShare, SAO and slime.

Current stories

build1 publisherOne report

NVIDIA's LoGRA fits RL post-training of a 27B model onto a single eight-GPU node

NVIDIA's LoGRA cuts average RL post-training memory by up to 45.7% by storing gradients as low-rank sketches, according to a write-up of its paper. With that saving, a 27B model trained stably for over 1,100 steps on one eight-GPU node where dense Adam runs out of memory.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence40