Build1 publisherNot yet confirmed elsewhere2 min readPublished
VIDRAFT ships a 111GB Darwin-180B that pages its experts from SSD on 8GB-VRAM laptops
VIDRAFT released a 111GB, 4-bit build of its 180B-parameter Darwin model that it says runs on a laptop with 8GB of VRAM and 32GB of RAM. Air-gapped teams still need a measured token rate on that laptop before they can price it against a server.
The Engineer · Build desk

What happened
- The model routes each token to 10 of its 512 experts, so compute per token corresponds to roughly 3B active parameters.
- llama.cpp streams weights from the SSD on demand, so only the expert weights for the current token have to be held in memory.
- VIDRAFT reports an unchanged, self-reported 87.65% MMLU-Pro score on 2,000 matched questions before and after quantization.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint With at least 71GB of weights on disk at all times, the laptop's SSD read speed and the router's cache hits set how fast it generates text.
- contradiction The post's own figures give 3.2x on storage and 60x on active compute, so its 250x hardware claim is unfit for sizing an on-prem deployment.
- exposure Adopters rely on one self-reported benchmark for the no-loss claim, and matching totals of 1,753 correct can hide questions lost and gained.
The laptop has 40GB of fast memory: 8GB on the GPU and 32GB of system RAM [11]. The GGUF build totals about 111GB [4]. So at least 71GB of weights sit on the SSD at any moment [12]. According to the dev.to post, llama.cpp streams weights from disk on demand, and only the expert weights the current token needs have to be resident [6]. The post says "the SSD becomes an extension of the memory hierarchy" [10]. For setup, it points readers to llama.cpp's mmap documentation and leaves the exact flags to the model card [9].
The router keeps the disk traffic bounded. Each token activates 10 of 512 experts, roughly 3B parameters of compute [5]. Storing 180B parameters [2] in 111GB works out to about 4.9 bits per weight [13], so the 4-bit label is generous by nearly a bit. At that density, 3B active parameters take up about 1.8GB [14]. One token can pull at most about that much from the SSD. Experts already in RAM from earlier tokens need no new read. The token rate therefore depends on how often the router picks experts outside the cache and how fast the drive returns them.
The post does not include a tokens-per-second figure for the 8GB laptop [1]. For air-gapped or on-prem work, speed decides the comparison with the multi-GPU server that the 360GB BF16 original conventionally needs [3].
The post also puts the hardware reduction at roughly 250x against that full-precision server [18]. Its own figures give 3.2x on storage, from 360GB to 111GB [15]. On active parameters, from 180B to 3B, the ratio is 60x [16]. Neither is near 250x.
The accuracy test is better designed than the headline ratio. VIDRAFT compared matched items: 87.65% on 2,000 MMLU-Pro questions before quantization and 87.65% after, a result the post labels self-reported [8]. Both runs scored 1,753 correct [17]. Equal totals can still hide questions the quantized model lost and others it gained. The post argues that the MoE design is why quantization loses less here than it would on a dense model [7]. For the score to carry over to an on-prem workload, that workload has to look like those 2,000 questions.
I think the engineering is sound. A sparse router over memory-mapped weights is a sensible way to fit 180B parameters into a 40GB machine [2] [11]. Whether it beats a server on cost depends on the drive inside the laptop.
What to watch
- An independent tokens-per-second measurement of POCKET-Darwin-180B on an 8GB-VRAM, 32GB-RAM laptop, with the SSD model stated.
- Independent replication of the 87.65% MMLU-Pro result, ideally with a per-question comparison of the BF16 and 4-bit runs.
- Whether VIDRAFT publishes how it calculated the roughly 250x hardware reduction.