Build1 publisher3 min readPublished
Strata Hits 110.8 tok/s on an RTX 4090 by Activating Only 6B Parameters per Token
Strata runs a 125B-class Qwen model at about 100 tok/s on one RTX 4090 because its sparse MoE activates only about 6B parameters per token. The design moves the hardware bill to system RAM, and the 100 tok/s figure holds mainly for code and structured output.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Strata keeps attention, shared weights and the most-used experts in the card's VRAM, parks the remaining experts in system RAM, and streams them over PCIe as tokens need them.
- An independent 4090 test on a Ryzen 7900 with 192 GB of DDR5 and a 280 W power cap measured 110.8 tok/s text decode at IQ3_S with low reasoning.
- The tester says every expert and the n-gram table stayed RAM-resident, the quality scores come from a single seed, and only one active request is optimized.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost The hardware bill moves from the GPU to system memory: the reference box carried eight times the card's 24 GB as DDR5, so the build that reproduces the number is a workstation.
- constraint With only one active request optimized, the measured speed describes a single developer's local session, and serving several users from one box is outside anything the tests cover.
- decision Low prose acceptance pulls chat-style use well under 100 tok/s, so the case for adopting it rests on agentic coding, where the model gains 16.5 DeepSWE points over the 27B.
- contradiction The SSD placement for the n-gram table comes from a different deployment; the headline 4090 run kept the table in RAM, so a lower-RAM version of this build has no 4090 figure behind it.
According to a dev.to write-up of the repo, the router sends each token to 10 of 512 experts plus one shared expert [3]. That is about 2% of the expert pool (10/512) [1]. Counting the n-gram embeddings and the prediction head, about 6B of roughly 180B parameters are touched per token, or about 3% [c1, d2]. A 24 GB card still cannot hold the weights [2]. Strata only needs the hot fraction on the card [4].
The published launch config shows how the hot set is chosen. `--expert-profile` points at a profile of which experts run hot, `--expert-cache auto` keeps those on the GPU, and `--pcie-frac 0.35` holds the bus below saturation [11]. PCIe bandwidth is the limit the design works around. Every expert the cache misses has to come over the bus from system RAM [4].
The n-gram table is the cleverest part of the stack. It holds about 51B parameters, roughly 28% of the total (51/180), and it is a lookup with sparse row reads. It does no matrix multiply [c5, d3]. Sparse reads can tolerate an SSD [5]. The write-up reports one deployment that kept a 47.7 GiB quantized table on SSD and still decoded at full speed [5]. The headline 4090 run did not use that layout. Its author says all experts and the n-gram table must stay RAM-resident, and the box had 192 GB of DDR5 [c8, c10]. A 4090 launched at $1,600, so "consumer hardware" was generous before the DDR5 went on the parts list [17].
Speculative decoding supplies the rest of the speed. A small multi-token-prediction head drafts tokens and the main model verifies them in one pass. The config drafts four ahead with `--spec 4 --spec-min-p 0.7` [c6, c11]. Acceptance runs about 0.59 on prose, 0.88 on code and 0.93 on structured output [6]. On a Strix Halo box, the same model decoded about 22 tok/s on prose and 82 tok/s on code [7]. That is a 3.7x spread from the workload alone (82/22) [4].
For the 110.8 tok/s figure to carry over, your setup has to match the measured one: IQ3_S quantization, low reasoning, a 280 W power cap, 192 GB of RAM and a single active request [c8, c10]. The quality evidence is thinner. The run scored 28 of 30 on an executable DevOps suite, from a single seed with one to two cases of variance [c9, c10]. Quantization is the other large knob. In Strata's own table for an RTX 5070, Q2_0 decoded 94 tok/s and IQ3_S decoded 53, about 1.8x apart (94/53) [c12, d5]. The write-up recommends IQ3_S as the default for output quality [12].
I think this build suits one developer running agentic coding locally, since one active request is all the tests optimize for [10]. The benchmarks point the same way. According to the write-up, Flash-Next beats Qwen3.8-27B on all nine shared benchmarks by an average of 4.19 points, and on DeepSWE it rises from 42.2 to 58.7 [13]. Anyone building on llama.cpp needs a build from Aug 27, 2026 or later. Older builds fail with `unknown model architecture: 'qwen4exp'` [14]. The installer detects GPU and RAM, picks a quant, downloads about 70 GB and serves an OpenAI-compatible API [15]. The license is the Qwen Community License 1.0, and it allows self-hosting and commercial use [16].
What to watch
- A 4090 measurement with the n-gram table on NVMe and well under 192 GB of RAM would show whether the decode speed survives on a smaller box.
- Multi-seed quality runs or multi-request throughput from the independent 4090 repo would test whether the 28/30 score and the single-user speed hold.
- A prose-only decode figure on the 4090 would show how far below 110.8 tok/s chat workloads land at 0.59 acceptance.