Jeff, an open 0.8B model, returns a probability for each option in one forward pass and sends low-confidence agent decisions to Qwen3.8-27B. The confidence score attached to each answer tells an agent when a small decision is worth the larger model's time.
Reality
- Evidence38
- Adoption15
- Hype gap+15
- Incentives
- Insufficient
- Confidence35
Polyglot's developer ran six coding agents 30 times on each of seven local models, and three never made a tool call on models that write calls as text. The author's own error bars say 30 runs can sort agents into tiers but cannot rank two agents inside one.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives70
- Confidence55
Researchers writing on LessWrong found only Nemotron 3 Super and Qwen3 32B changed refusals when test suspicion was removed from their reasoning traces. In most other models, test talk in a trace looks like general caution, so it is weak evidence of eval gaming.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence45
Portable Computer's harness is real engineering. The bill of materials starts at a $4,800 DGX Spark or a 24GB RTX card, which turns local agents into a hardware order.
Perspective Coverage
3 publishers
- Builder
- Builder 33%
- Operator
- Operator 40%
- Investor
- Investor 27%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives70
- Confidence60
LM Studio 0.4.0 shipped llmster, a headless daemon that runs it on the Linux GPU servers where MIT-licensed Ollama already worked. Teams whose policy demands auditable source now decide on licence, since only LM Studio's lms CLI carries an MIT grant.
Reality
- Evidence45
- Adoption35
- Hype gap+10
- Incentives20
- Confidence45
The startup says Anchor 3.0 caught more than 90 percent of violations in a test it published itself, at under a 500th of the cost of a frontier call, and the messages it misses stay the firm's problem.
Reality
- Evidence30
- Adoption16
- Hype gap+32
- Incentives82
- Confidence58
Alibaba's Qwen3.8-Flash-Next preview activates 6B of its 125B parameters per token. Per-token compute drops to under a quarter of the dense 27B's, and about 125GB of weights still has to stay on device.
Reality
- Evidence34
- Adoption18
- Hype gap+15
- Incentives60
- Confidence42
Bonsai 2 27B keeps 98 percent of Qwen3.8 27B's aggregate benchmark score, up from 95 percent for PrismML's first release, and TechCrunch reports the 5.9GB file is small enough for a PC and possibly a high-end phone.
Reality
- Evidence44
- Adoption38
- Hype gap+33
- Incentives71
- Confidence55
PrismML's ternary build of Qwen3.8 27B keeps 98.2 percent of the full-precision benchmark average on both of the company's inconsistent scorecards, and the loss it does take is concentrated in knowledge and reasoning.
Reality
- Evidence38
- Adoption20
- Hype gap+32
- Incentives72
- Confidence55
The calculator prices its featured coding workload at about $6.94 a month through OpenRouter against 26 cents of electricity. The machine is still $3,419 down after a year and 7.97 billion tokens from break-even.
Reality
- Evidence55
- Adoption15
- Hype gap+10
- Incentives30
- Confidence55
Every task in the benchmark is checked against its required terminal state, so any tool order can pass and any extra side effect fails. Claude Opus 5, the strongest model tested, cleared 66.50% at pass@1.
Reality
- Evidence60
- Adoption18
- Hype gap+12
- Incentives35
- Confidence52
A 15M-parameter model streams English text on a 2007 PSP at about one token per second. That is the extreme end of a sizing rule. The harder half of that rule is checking whether the file that fits is a format its own maintainer recommends.
Reality
- Evidence38
- Adoption31
- Hype gap+12
- Incentives58
- Confidence46
Nemotron 3.5 Lightning touches a tenth of its weights per token, which is what makes an agent loop plausible next to the sensors, but memory still has to hold all thirty billion and NVIDIA's own loop still escalates off-device.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+38
- Incentives84
- Confidence54
Prefill and decode sit on the same card and answer to different limits, which is why a GPU with more arithmetic and the same memory read rate leaves your time per output token exactly where it was.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+16
- Incentives28
- Confidence56
Meta's Muse Glimmer 30B and Alibaba's Qwen3.8-27B both landed in August under pure Apache 2.0 and both fit one 24 GB GPU, so the deployment question moves off licence terms and onto how you spend the memory that is left.
Reality
- Evidence16
- Adoption12
- Hype gap+52
- Incentives62
- Confidence20
JetBrains has taken the assembly work out of running a coding agent offline. What it could not take out is the hardware, and that is now the thing deciding who adopts.
Reality
- Evidence56
- Adoption18
- Hype gap+16
- Incentives72
- Confidence63
A custom Qwen 3.5 2B adapter returned a shell command when OpenCode's prompt said September 1st, 2026. The harness was running with auto-approval, and the numbers are one researcher's.
Reality
- Evidence44
- Adoption18
- Hype gap+8
- Incentives44
- Confidence52
The claim reaches software vendors through a syndicated summary carrying two incompatible hardware descriptions and no released artifacts. The weights, meanwhile, are Apache 2.0.
Reality
- Evidence22
- Adoption28
- Hype gap+48
- Incentives58
- Confidence44
A published NVFP4 and speculative-decoding config turns a 27B open-weights model into something you can try to serve. The 206.1 tokens per second figure is single-stream and unreplicated.
Reality
- Evidence52
- Adoption28
- Hype gap+14
- Incentives70
- Confidence46
Dynamic 3.0 ships Qwen3.8-27B GGUFs from 6.2GB up, with an unreproduced accuracy claim attached. The number that matters is the one that decides where the file fits.
Reality
- Evidence34
- Adoption45
- Hype gap+28
- Incentives74
- Confidence41