Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
An RTX 3090 holding transcription and embedding models in VRAM averaged 25 W over the month at a cost of 2 euros, while the same post puts hardware amortisation at about 25 euros a month, twelve times the power bill.
Reality
- Evidence35
- Adoption10
- Hype gap+25
- Incentives50
- Confidence40
nano-vLLM's decode cost fits in one expression, W/B plus KV(Tavg), and a batch-1 context sweep on Qwen3-0.6B shows why those two terms move in opposite directions as you add requests or lengthen prompts.
Reality
- Evidence47
- Adoption
- Insufficient
- Hype gap+12
- Incentives25
- Confidence45
A residency policy that barred even embedding calls from leaving the building forced a full local RAG stack. The hardware math turns out to be the easy part.
Reality
- Evidence28
- Adoption18
- Hype gap+32
- Incentives46
- Confidence33
A week-long failure log on two RTX 3090s under WSL2 lands on one config at 170-210 tok/s. Everything before it died in dependency resolution, not in the math.
Reality
- Evidence38
- Adoption18
- Hype gap−12
- Incentives27
- Confidence44