One developer's test on 484 SEPA rulebook passages found that plain-English questions push several answers out of a top-5 vector search. Because the test measures each answer's rank directly, the failure shows up in retrieval, before the language model writes anything.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence45
Deferred tool loading assumes retrieval puts the right tool in the shortlist. A BM25 harness over 100 synthetic enterprise tools and 200 tasks says it does that 5 percent of the time when the user does not use your words.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives25
- Confidence60
Name-only BM25 matched nothing for "money we gave back to shoppers". The description channel put the right table second and embeddings put it twenty-second, and reciprocal rank fusion let the blind channel cost the answer nothing.
Reality
- Evidence64
- Adoption15
- Hype gap−10
- Incentives60
- Confidence58
The loader read a Spider 2.0 checkout with rglob and swallowed every read error, so 2,868 files Windows would not open became tables that do not exist, and the benchmark scored recall against what was left.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+5
- Incentives25
- Confidence55
A dev.to post argues that IBM's chunkless RAG only helps on documents whose structure a parser can actually recover, and that on scanned PDFs and one-table wikis the agent ends up walking a tree the parser invented.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+34
- Incentives44
- Confidence36
myc is an MIT-licensed memory layer for coding agents that installs into Claude Code's PreCompact event, writes the raw session to disk before it prints anything back, and hides the decisions it extracts until a human confirms them.
Reality
- Evidence38
- Adoption10
- Hype gap+12
- Incentives65
- Confidence45
Quantization pins a compressed copy in RAM and moves the float32 originals to disk, so every rescore becomes a disk read. Choose a smaller datatype instead and rescore has only the smaller vectors to score against.
Publishers:qdrant.tech
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−10
- Incentives70
- Confidence55
A CQADupStack benchmark reports dense retrieval beating BM25 by more on identifier-bearing queries than on the rest, because the identifier is often absent from the documents that answer them.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+12
- Incentives42
- Confidence54