Build1 publisher3 min readPublished
Ternary Bonsai 2 27B runs at two to three tokens a second on a GPU-less VPS
Prism's ternary Bonsai 2 27B ran at two to three tokens a second on a CPU-only Hetzner VPS in a dev.to test, against Simon Willison's 20 to 44 on a Mac. The sub-6GB file fits a 16GB box easily, but at that speed it only suits batch jobs.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Prism's Ternary-Bonsai-2-27B-gguf packs a 27-billion-parameter model into a single GGUF file just under 6GB, stored in a format called PTQ1_0.
- Most mainstream llama.cpp builds cannot read PTQ1_0 because its 1-bit support is not merged upstream, so the model needs Prism's own llama.cpp fork.
- The download took five and a half minutes on the VPS, for a model that would otherwise arrive as more than 40GB uncompressed.
- Once the context window and server overhead are added, the author said to plan for 8 to 9GB of RAM; the test VPS had 16GB.
- Restarting the server sometimes changed throughput by close to 40 percent with no configuration change, an oddity Willison had also seen.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint On a CPU-only host the model fits queued, unattended work; any interface where a person waits on the reply needs Metal or CUDA hardware.
- cost Running it means tracking a vendor fork of llama.cpp until PTQ1_0 lands upstream, so runtime updates depend on Prism's release schedule.
- decision An 8 to 9GB working set at an 8,192-token context puts 8GB instances out of bounds even though the model file is under 6GB.
Under 6GB for 27 billion parameters averages out to less than 1.8 bits per weight [1]. The format gets there by collapsing every weight to one of three values, roughly minus one, zero or plus one, and relying on scaling factors to recover something close to the original behavior [2]. That encoding fixes the file size. The file is the same on a Mac and on a VPS, so the "9x smaller" half of the release pitch [1] holds on cheap hardware by construction.
Speed does not carry over. Willison's run used the Apple Silicon build with Metal [6], and it was roughly 7 to 22 times faster than the CPU-only VPS [2]. The author wrote: "Usable for a background batch job. Not usable for anything where a person is sitting there waiting." [8]
For the VPS number to predict another CPU host, that host would need similar cores and memory. The author describes the machine only as the small Hetzner VPS that runs the blog's publishing pipeline, with 16GB of RAM and no GPU [16]. Restarts moved throughput by as much as 40 percent [10], so I would average several launches before sizing anything on this model.
The memory estimate comes from one config line: `llama-server -m bonsai.gguf --port 8331 -c 8192` [12]. The 8,192-token context is part of the 8 to 9GB figure [11], so a longer context raises it. On the 16GB box, the author's estimate leaves 7 to 8GB free [3]. The author also advises running `free -h` before starting the server, not after it dies [15].
The packaging is good work. Prism shipped CPU-only Ubuntu binaries next to the CUDA and Metal builds [5]. As a result the whole setup was two curl downloads, a tar extract and the server launch [12]. The catch is the pin. The runtime is build prism-b10685-7dffb15 of a fork [12]. I'd count that as the main adoption cost, because until 1-bit support is merged upstream, llama.cpp fixes reach this runtime only when Prism rebases [4].
The evidence on quality is thinnest. "I did not run a proper benchmark suite against the full-precision version of this model," the author wrote [13]. The test used everyday jobs instead: summarizing documentation, writing a short regex, explaining a stack trace and renaming variables. On those, the model "held up better than I expected" [14]. For that result to apply to another workload, the prompts would have to look like those four, and they would need scoring against a full-precision baseline. In my view the 9x figure holds on commodity hardware. Near-lossless is still the release's claim, checked here against one person's four task types [1] [14].
What to watch
- Upstream llama.cpp merging PTQ1_0 support, which would remove the dependency on Prism's fork.
- A benchmark suite scoring Ternary Bonsai 2 27B against its full-precision base, which would test the near-lossless claim directly.
- CPU-only throughput reports that state core count and CPU model and average across several server restarts.