Strata runs a 125B-class Qwen model at about 100 tok/s on one RTX 4090 because its sparse MoE activates only about 6B parameters per token. The design moves the hardware bill to system RAM, and the 100 tok/s figure holds mainly for code and structured output.
Reality
- Evidence45
- Adoption15
- Hype gap+40
- Incentives
- Insufficient
- Confidence40
Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Mert Cobanov's open-source Ollaya serves local decision models behind a Jev-compatible API and reports 8 to 10 ms per request on an RTX 4090. Its comparisons with TypeSafe's hosted service are uncontrolled, so teams have to check accuracy on their own traffic before switching.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence35
OpenBMB's 2B-parameter model is Apache-2.0, speaks 30 languages, and serves through vLLM's OpenAI-compatible /v1/audio/speech. The real-time factor quoted for it was measured on a different backend than the one the project recommends for production.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+38
- Incentives42
- Confidence40
The sample times the GPU decode correctly with CUDA events, then adds a Tier-2 parse term that a cast to whole seconds rounds to zero on every frame. Fastvideo puts that missing stage at 15 to 29% of decode time.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+14
- Incentives78
- Confidence55
After weeks of reading 100 percent on its own curriculum, one team rebuilt its coding-agent bench around unseen prompts and guard redirects per run, and found two harness contaminations on the way.
Reality
- Evidence36
- Adoption11
- Hype gap−9
- Incentives58
- Confidence42
A residency policy that barred even embedding calls from leaving the building forced a full local RAG stack. The hardware math turns out to be the easy part.
Reality
- Evidence28
- Adoption18
- Hype gap+32
- Incentives46
- Confidence33
A dev.to guide argues local LLM capacity planning collapses into one napkin equation. Run it first and the hardware shortlist writes itself, tier names and all.
Reality
- Evidence42
- Adoption20
- Hype gap+24
- Incentives55
- Confidence38