Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

NVIDIA's LoGRA fits RL post-training of a 27B model onto a single eight-GPU node

NVIDIA's LoGRA cuts average RL post-training memory by up to 45.7% by storing gradients as low-rank sketches, according to a write-up of its paper. With that saving, a 27B model trained stably for over 1,100 steps on one eight-GPU node where dense Adam runs out of memory.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying NVIDIA's LoGRA fits RL post-training of a 27B model onto a single eight-GPU node
Photo: nvidia.com

What happened

  • LoGRA stores each compressed gradient as a random low-rank sketch, shrinking a d-by-k gradient buffer to d-by-r, a reduction by a factor of k/r.
  • A custom optimizer, RowAdam, keeps running estimates of squared row magnitudes and drops the first-moment momentum that Adam carries.
  • Before each update, a predicted-KL controller estimates how far the next-token distribution would shift and scales the step to keep that shift within a set budget.
  • The authors tested Qwen models from 1.5B to 27B parameters on mathematical reasoning benchmarks.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability RL post-training at 27B comes within reach of a team that owns one eight-GPU node, a hardware budget the write-up says left serious RL runs out of reach for most practitioners.
  • decision Adopting LoGRA swaps exact gradients for a rank-r estimate. Training stability then depends on the KL budget each team chooses and has to defend.
  • cost The controller shrinks steps whenever the sketch predicts a large policy shift, so some of the memory saving may be repaid in extra training steps.

The evidence here is a dev.to write-up of the paper from NVIDIA and collaborators [3]. According to that write-up, gradients and optimizer states in RL post-training can take several times the memory of the model weights [15]. Under dense Adam, a 7B model's gradient and optimizer buffers alone can push peak memory past 30 GiB per GPU. That is before activations or the rollout buffer are counted [12]. FSDP and CPU offloading spread that state across devices or move it to host memory. The gradient tensors stay the same size [13].

I think LoGRA cuts the right buffer. During backpropagation it accumulates S = G * A^T, where G is the d-by-k gradient and A is a random r-by-k matrix with r much smaller than both d and k [5]. The full-size estimate, S * A, exists only at the moment the update is applied [6]. Only attention and MLP projection matrices get this treatment, and the write-up says those hold the bulk of gradient memory [8]. Nodes exchange the same compact sketch for policy synchronization, so the saving extends to communication [7].

RowAdam is the piece of good engineering in the design. Adam keeps first and second moment estimates for every parameter [14]. A per-parameter first moment would put a full d-by-k buffer back beside every compressed matrix. Dropping it and tracking a statistic per row, as RowAdam does, keeps optimizer state at row scale [9].

The KL controller is there because compression creates a new way to fail. An approximate gradient can produce an oversized update, and in RL a large policy jump can derail training [16]. According to the write-up, the KL prediction is computed from the sketch itself, with no separate forward pass [11]. The 27B run is the evidence that the controller holds at that scale, under whatever budget the authors set [4].

The 45.7% figure is the best case for average training memory saved, and the write-up says task performance held [1]. At that best case, a run still needs 54.3% of its dense-Adam memory [18]. The available text of the write-up stops before the per-model memory table and does not name the GPU in the eight-GPU node. For the single-node 27B result to carry over to another team, that team's cards need at least the authors' per-GPU memory. Its task also has to look like the math reasoning benchmarks the tests used [2].

What to watch

  • Wall-clock time and steps-to-target-reward for LoGRA against dense Adam at a size where both fit, such as 7B, to show what the conservative KL steps cost in training time.
  • Results from LoGRA on RL tasks outside the math reasoning benchmarks, such as code or agent work, where the reward signal behaves differently.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories