Verbalization training made three models voice test suspicion 2.4 to 2.9 times as often in chain of thought, with task behavior largely unchanged. For monitors, it suggests how often a model reports a belief can be trained apart from the belief itself.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence40
PAIR proxies Ollama and LM Studio, so the agent keeps seeing one connection and no harness code changes. The adoption cost moves to disk, because a node is only eligible if it already has the exact model downloaded.
Perspective Coverage
7 publishers
- Builder
- Builder 45%
- Operator
- Operator 38%
- Investor
- Investor 17%
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+30
- Incentives72
- Confidence60
An instrumented run from pod creation to first response found eight minutes spread over six phases, with kernel recompilation eating a 64 GB model's startup and an S3 download pattern eating a 203 GB model's.
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap+28
- Incentives76
- Confidence43
FreeToken's preprint puts DeepSeek-V4-Flash on a single RTX 5090 desktop. The GPU holds about a sixth of the machine's memory, and nobody outside the team has run the benchmark yet.
Reality
- Evidence44
- Adoption12
- Hype gap+16
- Incentives58
- Confidence52
A single-box test in Japan put 76 tokens/s next to 4.4 tokens/s, then found the deciding variable elsewhere: which models held the output format and which invented reassurance.
Reality
- Evidence52
- Adoption22
- Hype gap−8
- Incentives45
- Confidence55