Jeff, an open 0.8B model, returns a probability for each option in one forward pass and sends low-confidence agent decisions to Qwen3.8-27B. The confidence score attached to each answer tells an agent when a small decision is worth the larger model's time.
Reality
- Evidence38
- Adoption15
- Hype gap+15
- Incentives
- Insufficient
- Confidence35
Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
Capsule Security has shipped an evaluator that runs inside an agent's execution path and can stop an intended action before it executes. Doing that requires adding a chokepoint to the agent's execution path.
Reality
- Evidence35
- Adoption15
- Hype gap+35
- Incentives75
- Confidence55
The startup says Anchor 3.0 caught more than 90 percent of violations in a test it published itself, at under a 500th of the cost of a frontier call, and the messages it misses stay the firm's problem.
Reality
- Evidence30
- Adoption16
- Hype gap+32
- Incentives82
- Confidence58
A case study reports an 8B model that got better at naming threat categories while losing ground on severity scoring, lifecycle counting and declining a weather question. Its author says every failure mode is already in the literature.
Reality
- Evidence25
- Adoption12
- Hype gap+8
- Incentives30
- Confidence55
A Microsoft case study credits Docusign's swap to small task-specific models with 90% lower cost and eight times the throughput. Docusign's own accounting of the whole pipeline claims 50 times cheaper per document.
Reality
- Evidence42
- Adoption62
- Hype gap+24
- Incentives80
- Confidence55
A KDnuggets walkthrough builds its comparison on Qwen2.5-0.5B-Instruct on an M2 MacBook Air, where the keys and values for a static instruction block come out identical on every call and only the ticket text changes.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives20
- Confidence55
ShadowPEFT is in Hugging Face PEFT's main branch as of a September 15th announcement, and it carries its own hidden state and can be detached as a smaller standalone model. Trying it means installing PEFT from source.
Reality
- Evidence58
- Adoption20
- Hype gap+10
- Incentives65
- Confidence57
Microsoft's local runtime handles model download, hardware detection and execution provider selection behind a .dll, .so or .dylib. An app that adopts it owns the download, the cache and the unload.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence40
The OpenAI bill it displaced was about EUR 10 a month. Paying that back takes a second set of calibration values, and the only comparison actually benchmarked in the writeup is embeddings.
Reality
- Evidence35
- Adoption12
- Hype gap+5
- Incentives30
- Confidence45
A Llama 3.2 3B agent running offline invented a $1,990 balance on a $1,975 invoice. The open-sourced answer leaves prose to the model and hands every number to deterministic Python behind a tri-state router.
Reality
- Evidence38
- Adoption10
- Hype gap+22
- Incentives58
- Confidence48
A 15M-parameter model streams English text on a 2007 PSP at about one token per second. That is the extreme end of a sizing rule. The harder half of that rule is checking whether the file that fits is a format its own maintainer recommends.
Reality
- Evidence38
- Adoption31
- Hype gap+12
- Incentives58
- Confidence46
Stanford's AI Index puts a 142-fold parameter cut and a more than 280-fold price cut behind one fixed MMLU threshold. That narrows where building your own still pays, and it lands in a state-law count that doubled in a year.
Publishers:hai.stanford.edu
Reality
- Evidence60
- Adoption68
- Hype gap+10
- Incentives40
- Confidence57
Microsoft Research says a 4-billion-parameter model trained with SocialRL out-scored GPT-4.1, GPT-5.1 and GPT-5.2 across six bargaining domains. The winning margin is thinner than the method's effect.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+46
- Incentives68
- Confidence33
A Fast Company column argues enterprise model selection is now cost optimisation. The consequence is that the layers deciding outcomes, context and feedback, are the buyer's own build.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+34
- Incentives38
- Confidence45
Alibaba's Apache-2.0 Qwen3.8-27B fits in about 17GB and matched near-frontier scores, per Artificial Analysis. It also burned 3.7x the median output tokens getting there.
Reality
- Evidence62
- Adoption64
- Hype gap+18
- Incentives60
- Confidence55