Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
An Apache-2.0 memory layer for coding agents returns a candidate only when several signals agree, and it labels every answer STRONG, WEAK or MISS so the calling agent has to branch on confidence before reusing an old fix.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence52
Open Walnut patched QMD's compiled output 15 times from the outside, then found that the one stage it needed to change, tokenization, lived in the engine core. Writing its own engine again took eleven days.
Reality
- Evidence60
- Adoption45
- Hype gap−10
- Incentives55
- Confidence55
The author of Doco argues that browse, search, cite and safe-edit operations over a maintained corpus do most of the work teams expect from embeddings, and that the pipeline adds a copy which drifts from the source.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+12
- Incentives68
- Confidence52
Langhuan's first retrieval benchmark found the keyword channel at 0.0000 recall and hybrid search matching vector-only digit for digit. The labeled data came free.
Reality
- Evidence61
- Adoption14
- Hype gap−6
- Incentives52
- Confidence55
A read-only GraphQL service over an existing SQLite file cut a four-request search screen down to one. The interesting parts are DataLoader, cost limits and a query_only pragma.
Reality
- Evidence38
- Adoption18
- Hype gap+6
- Incentives32
- Confidence44
Six new feature areas landed in v3.0.0, every one off by default, and a verify command that refuses to call an unreadable database clean.
Reality
- Evidence48
- Adoption12
- Hype gap−8
- Incentives66
- Confidence46