A KDnuggets walkthrough builds its comparison on Qwen2.5-0.5B-Instruct on an M2 MacBook Air, where the keys and values for a static instruction block come out identical on every call and only the ticket text changes.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives20
- Confidence55
The coding-trained model came out ahead over thirty runs, but four of the five fixtures tied, which leaves the whole margin in a single Docker log task run at default temperature between a 30b model and an 8b one.
Reality
- Evidence58
- Adoption10
- Hype gap−15
- Incentives25
- Confidence55
A 9B MIT-licensed coding model that reportedly matches a 31B rival on SWE-Bench Verified is still unusable as a Claude Code backend, because the runtime never turns its tool-call XML into a file write.
Reality
- Evidence44
- Adoption18
- Hype gap+38
- Incentives24
- Confidence52
p-for-llm puts a 29-expert, top-1-routed mixture of experts on an ESP32-P4 and got two points on Hacker News. The routing, not the quantization, is the load-bearing part.
Reality
- Evidence28
- Adoption8
- Hype gap+18
- Incentives45
- Confidence34
A 28.9-million-parameter model now runs on an ESP32-S3 that was topping out at 260K two years ago. The gain came from moving most of the weights into flash and reading 450 bytes per token.
Reality
- Evidence55
- Adoption24
- Hype gap+20
- Incentives32
- Confidence47