Skip to content

Build1 publisher3 min readPublished

Your GPU reports 24GB. Only 7.9GB of it loads a model, and half of that is already gone

A single 8GB laptop produced four different VRAM figures, none of them a bug. The one that decides whether a model loads is the one no tool puts in front of you.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The machine described is a laptop with an RTX 5060 and 8GB of VRAM.
  • Running dxdiag /whql:off /t on the machine and opening the Display tab reports the graphics card as having 24,144 MB of display memory.
  • Only the dedicated memory figure decides whether a model loads.
  • The dxdiag Display tab breaks the total down as Display Memory 24144 MB, Dedicated Memory 7899 MB, Shared Memory 16245 MB, and 7,899 plus 16,245 equals 24,144.
  • Display memory is the dedicated VRAM on the card plus a slice of system RAM the driver treats as overflow, half the machine's 31.7GB, the standard Windows allocation.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

On an RTX 5060 laptop sold with 8GB of VRAM, dxdiag reports the graphics card as having 24,144 MB of display memory [1][2]. Nothing is broken: the figure is arithmetically honest, and it is also not the number that decides whether a model loads [3][5].

Four lines down the same Display tab, the panel breaks it out: 7,899 MB dedicated, 16,245 MB shared, 24,144 MB total [4]. The shared pool is a slice of system RAM the driver treats as overflow, half the machine's 31.7GB, which is the standard Windows allocation [5]. Only the dedicated figure governs whether a model fits in fast memory [3]. Taking the headline number at face value overstates usable capacity by roughly 3.1x on this machine [4].

It gets worse if you add adapters up. Every display device on the box claims the same 16,245 MB: integrated Intel graphics, the RTX 5060, and two DisplayLink devices [6]. Sum what the machine appears to have and you get 64,980 MB of shared memory that exists exactly once [1]. The dedicated figure is not perfectly stable either. A model-fitting tool on the same box reported 7.96GB against dxdiag's 7,899 MB [7], a gap of about 61 MB [5], close enough to be the same answer and far enough apart that a tight fit calculation will disagree with itself depending on which tool it asked [7].

The two pools behave differently under load, which is the part that costs you. Spilling into shared memory does not fail, it slows down [8]. A 9B model at 32K context sits entirely in VRAM at 5.9GB and generates 51.5 tokens/sec; push it to 64K, it needs 7.5GB, splits 18/82, and drops to 22.6 tokens/sec [9][10]. Same model, same machine, same quantisation, one setting, and throughput falls to about 44% of what it was [2]. A failure tells you something is wrong. A slowdown lets you keep going, wondering why everything feels off [8].

Then subtract the desktop. Under ordinary use, browser and editor open, between 2 and 3.5GB of the dedicated 7.9GB is spoken for before a model loads, and not in a way you reclaim by closing one app [11]. That leaves an effective budget of roughly 4.4 to 5.9GB [3]. The spread is the problem: measured with a browser, a notes app, a spreadsheet and around thirty other things touching the GPU it sat at almost exactly 2GB, and on a quiet afternoon it was under 1GB [12][13].

That is why the published benchmark in this account carries a footnote rather than a headline. A 36B mixture-of-experts model was measured three ways at 32K context, same day, flash attention on, KV cache quantised to q8_0, with ambient GPU load between 0.3 and 0.8GB against a normal desktop load of around 3.5GB [14][15]. The fastest row leaves about 800 MB of headroom and does not fit at all under normal desktop load [16]. The best number in the table cannot be reproduced while actually using the computer [16].

The rule that came out of it for dense models: about 7GB or smaller at 4-bit, and it has to run entirely on the GPU, because what spilling costs is more than a bigger model would have gained [17].

Worth watching: measure your own ambient GPU load on a busy day, not a quiet one, before trusting any fit calculator, and treat any published headroom figure that does not state the ambient load it was measured under as unreproducible.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories