Build1 publisher2 min readPublished Updated
VIDRAFT's 4-bit Darwin-180B build streams most of its weights from SSD on a 40 GB laptop
VIDRAFT says POCKET-Darwin-180B, a 111 GB 4-bit GGUF build of its 180B mixture-of-experts model, runs in llama.cpp on about $1,400 of consumer hardware. The accuracy evidence so far is one MMLU-Pro comparison that VIDRAFT reports itself.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Each token is routed through 10 of the model's 512 experts, so only about 3 billion parameters are active per token, according to a dev.to write-up of the release.
- Experts not held in memory are streamed from SSD on demand through memory mapping, and the write-up names storage I/O bandwidth as the bottleneck.
- The stated targets are a laptop with as little as 8 GB of VRAM and 32 GB of system RAM, or a CPU-only mini PC with about 128 GB of unified RAM.
- The BF16 original took 360 GB across 131 files, and VIDRAFT credits a method it calls graft quantization for the four-file build.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Anyone pricing the laptop build has to treat sustained SSD read bandwidth as the spec that sets tokens per second, ahead of the choice of CPU.
- constraint A speed result from the 128 GB mini PC will not predict the laptop, because one holds the whole file in memory and the other pages most of it from disk.
- cost Task-level accuracy checks fall to each adopter, since parity with the BF16 original has been shown on one aggregate benchmark only.
- capability If the claims hold, a stock llama.cpp build is enough to evaluate a 180B-class model on a single local machine with no custom tooling.
Ten experts out of 512 is about 2% of the pool on any one token [1]. That small slice is why SSD-backed inference can work at all, and the slice moves from token to token. If the active weights are stored at the file's average density, each token touches roughly 1.85 GB of them [2].
The laptop setup has 40 GB of fast memory. That leaves 71 GB of the 111 GB file, about 64%, on the drive [3]. Assume the router spreads load evenly across experts and the page cache holds nothing beyond what fits. Then about 1.2 GB of each token's weights comes off the SSD [4], and every gigabyte per second of sustained read buys a little under one token per second [5]. Real throughput beats that floor only as far as the router reuses the same experts on consecutive tokens.
The 128 GB mini PC is a different case. The whole file fits with about 17 GB to spare [6], so once the pages are warm nothing streams from disk and memory bandwidth sets the pace. A dev.to write-up of the release prices a "consumer setup" at about $1,400 [4] without saying which of the two machines that figure describes.
The file size also says something about the format. A flat 4 bits across 180 billion weights would come to 90 GB, and 111 GB works out to about 4.9 bits per weight [7]. Either some tensors are held above 4 bits or the format carries scale overhead. VIDRAFT reports zero accuracy loss on public benchmarks, according to the post [9]. Its author guesses at a targeted or layer-aware scheme and notes that standard Q4 quantization usually costs measurable quality [9]. For the MMLU-Pro parity [10] to transfer to another workload, that workload has to tolerate quantization error as well as one aggregate benchmark score does.
In my view the design is good engineering for the constraint it targets. Sparse routing plus memory mapping lets a file nearly three times the laptop's fast memory run at all [3]. The 250-to-1 cost ratio against a roughly $350,000 deployment of four to eight enterprise GPUs [4][8] sets a machine bought for capacity against one bought for speed. VIDRAFT positions the build for local development and evaluation [15]. To get the files, the post tells readers to search Hugging Face for VIDRAFT/POCKET-Darwin-180B [13]. For a release about leaving the data center, it also points readers to n1n.ai, a hosted OpenAI-compatible gateway, for API access [14].
What to watch
- Independent llama.cpp runs on the 8 GB VRAM, 32 GB RAM laptop configuration that report tokens per second and name the SSD used.
- Task-level comparisons of the 4-bit build against the BF16 original beyond MMLU-Pro, on coding, long-context or tool-use work.
- A published Hugging Face repository for the GGUF files, whose per-tensor quantization types would show what graft quantization does.