Build1 publisherNot yet confirmed elsewhere2 min readPublished
NVIDIA's LoGRA fits RL post-training of a 27B model onto a single eight-GPU node
NVIDIA's LoGRA cuts average RL post-training memory by up to 45.7% by storing gradients as low-rank sketches, according to a write-up of its paper. With that saving, a 27B model trained stably for over 1,100 steps on one eight-GPU node where dense Adam runs out of memory.
The Engineer · Build desk

What happened
- LoGRA stores each compressed gradient as a random low-rank sketch, shrinking a d-by-k gradient buffer to d-by-r, a reduction by a factor of k/r.
- A custom optimizer, RowAdam, keeps running estimates of squared row magnitudes and drops the first-moment momentum that Adam carries.
- Before each update, a predicted-KL controller estimates how far the next-token distribution would shift and scales the step to keep that shift within a set budget.
- The authors tested Qwen models from 1.5B to 27B parameters on mathematical reasoning benchmarks.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability RL post-training at 27B comes within reach of a team that owns one eight-GPU node, a hardware budget the write-up says left serious RL runs out of reach for most practitioners.
- decision Adopting LoGRA swaps exact gradients for a rank-r estimate. Training stability then depends on the KL budget each team chooses and has to defend.
- cost The controller shrinks steps whenever the sketch predicts a large policy shift, so some of the memory saving may be repaid in extra training steps.
The evidence here is a dev.to write-up of the paper from NVIDIA and collaborators [3]. According to that write-up, gradients and optimizer states in RL post-training can take several times the memory of the model weights [15]. Under dense Adam, a 7B model's gradient and optimizer buffers alone can push peak memory past 30 GiB per GPU. That is before activations or the rollout buffer are counted [12]. FSDP and CPU offloading spread that state across devices or move it to host memory. The gradient tensors stay the same size [13].
I think LoGRA cuts the right buffer. During backpropagation it accumulates S = G * A^T, where G is the d-by-k gradient and A is a random r-by-k matrix with r much smaller than both d and k [5]. The full-size estimate, S * A, exists only at the moment the update is applied [6]. Only attention and MLP projection matrices get this treatment, and the write-up says those hold the bulk of gradient memory [8]. Nodes exchange the same compact sketch for policy synchronization, so the saving extends to communication [7].
RowAdam is the piece of good engineering in the design. Adam keeps first and second moment estimates for every parameter [14]. A per-parameter first moment would put a full d-by-k buffer back beside every compressed matrix. Dropping it and tracking a statistic per row, as RowAdam does, keeps optimizer state at row scale [9].
The KL controller is there because compression creates a new way to fail. An approximate gradient can produce an oversized update, and in RL a large policy jump can derail training [16]. According to the write-up, the KL prediction is computed from the sketch itself, with no separate forward pass [11]. The 27B run is the evidence that the controller holds at that scale, under whatever budget the authors set [4].
The 45.7% figure is the best case for average training memory saved, and the write-up says task performance held [1]. At that best case, a run still needs 54.3% of its dense-Adam memory [18]. The available text of the write-up stops before the per-model memory table and does not name the GPU in the eight-GPU node. For the single-node 27B result to carry over to another team, that team's cards need at least the authors' per-GPU memory. Its task also has to look like the math reasoning benchmarks the tests used [2].
What to watch
- Wall-clock time and steps-to-target-reward for LoGRA against dense Adam at a size where both fit, such as 7B, to show what the conservative KL steps cost in training time.
- Results from LoGRA on RL tasks outside the math reasoning benchmarks, such as code or agent work, where the reward signal behaves differently.