Ternary weights at 1.76 bits put a 27B model into a 5.9GB file and let a laptop decode it at 28.1 tokens a second. The retention figure comes from Prism ML's own benchmark suite, not the table on the model card.
Reality
- Evidence38
- Adoption42
- Hype gap+34
- Incentives74
- Confidence46
Apple's M5 Ultra Mac Studio starts at $5,499 and reached $12,299 as tested, and the reviewer running local agents on it every day says the most expensive Anthropic subscription would still have cost less for years.
Perspective Coverage
10 publishers
- Builder
- Builder 36%
- Operator
- Operator 38%
- Investor
- Investor 26%
Reality
- Evidence82
- Adoption41
- Hype gap+18
- Incentives58
- Confidence76
Hugo Vergnes reports 0.384 CORE from 65.3 billion tokens on eight rented B200s. The $998 buys 43 hours of node time, not the 139 GPU-hours of failed run that produced the recipe, and not the evenings that wrote the framework.
Reality
- Evidence52
- Adoption12
- Hype gap+22
- Incentives45
- Confidence62
Alibaba's 27B model fits a 32GB card with 15GB to spare. Tom's Hardware still had to pick between llama.cpp's full 262K window at half-hour prefill and a supported vLLM deployment capped at 32K.
Reality
- Evidence60
- Adoption38
- Hype gap+15
- Incentives55
- Confidence58
NVIDIA Research rebuilt the serving stack for someone else's open-weight video model and got a five-second clip out in 1.653 seconds, but the default profile is lossy and the dense comparison run is the honest baseline.
Reality
- Evidence62
- Adoption45
- Hype gap+30
- Incentives72
- Confidence58
PAIR cut a five-subagent inbox task from 18 minutes to 8 minutes 48 seconds across three machines, which is just over two times the speed for three times the hardware, and every box in NVIDIA's demo cluster was one NVIDIA sells.
Reality
- Evidence32
- Adoption15
- Hype gap+25
- Incentives72
- Confidence45
Mithil Vakde's from-scratch transformer landed one point behind TRM on the public eval. His own ablations put 20 of those 44 points on two representation choices rather than on any amount of compute.
Reality
- Evidence44
- Adoption18
- Hype gap+16
- Incentives72
- Confidence41
JetBrains has taken the assembly work out of running a coding agent offline. What it could not take out is the hardware, and that is now the thing deciding who adopts.
Reality
- Evidence56
- Adoption18
- Hype gap+16
- Incentives72
- Confidence63
A review of 131 Reddit threads puts the deciding costs off the rate card: transfer bills that beat the compute bill, and storage pinned to regions where the GPU you need will not start.
Reality
- Evidence31
- Adoption
- Insufficient
- Hype gap+18
- Incentives62
- Confidence42
A dev.to guide argues local LLM capacity planning collapses into one napkin equation. Run it first and the hardware shortlist writes itself, tier names and all.
Reality
- Evidence42
- Adoption20
- Hype gap+24
- Incentives55
- Confidence38