NVIDIA's LoGRA fits RL post-training of a 27B model onto a single eight-GPU node
NVIDIA's LoGRA cuts average RL post-training memory by up to 45.7% by storing gradients as low-rank sketches, according to a write-up of its paper. With that saving, a 27B model trained stably for over 1,100 steps on one eight-GPU node where dense Adam runs out of memory.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40