Databricks made Lakebase Search generally available, pairing BM25 with a Postgres vector index it says runs 4 times cheaper than pgvector at 100M vectors. The vector index lives in object storage behind a cache, so large agent corpora no longer need RAM sized to the whole index.
Publishers:databricks.com · neon.tech Reality
- Evidence40
- Adoption25
- Hype gap+35
- Incentives90
- Confidence50
LiveReview's Maneshwar says Gemini File Search cost over 50 cents a review run because a reasoning model performed each search. A local index and a cheaper model brought runs down to 4 cents.
Reality
- Evidence45
- Adoption10
- Hype gap+15
- Incentives35
- Confidence40
A two-stage eval forces every shortlist to hold the correct tool plus its four strongest BM25 siblings. Selection accuracy comes in at 90 to 97 percent. That puts the 5 percent paraphrase recall on the critical path.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives35
- Confidence55
Deferred tool loading assumes retrieval puts the right tool in the shortlist. A BM25 harness over 100 synthetic enterprise tools and 200 tasks says it does that 5 percent of the time when the user does not use your words.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives25
- Confidence60
A Magic: The Gathering rules agent works out which source is entitled to settle a question before it answers. Its evaluation pits 496 structured documents against a single BM25 pass over the same corpus, graded blind.
Reality
- Evidence45
- Adoption12
- Hype gap−10
- Incentives65
- Confidence50
The build log for a Python package finder shows a day-one no-key rule pushing the download data onto a public JSON file and 15,000 PyPI API calls per refresh, held to a few minutes by a semaphore of 25.
Reality
- Evidence42
- Adoption10
- Hype gap+12
- Incentives55
- Confidence45
An embedding swap to bge-large made two long-failing negative tests pass because the retriever stopped surfacing the trap chunks at all. A chunk-id check in CI gives the suite a third verdict, test did not run.
Reality
- Evidence45
- Adoption12
- Hype gap+20
- Incentives65
- Confidence50
The extension puts BM25 ranking behind an ordinary Postgres index, so a newly committed row is searchable without a trip to an Elasticsearch cluster. PlanetScale's own benchmark reports 199 queries per second against ParadeDB's 7.9.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+31
- Incentives82
- Confidence57
The four RAGAS-lineage metrics were built to separate a retriever's failures from a generator's. Read in pairs, they also expose the case where the model skipped the context and got the answer right anyway.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence35
Name-only BM25 matched nothing for "money we gave back to shoppers". The description channel put the right table second and embeddings put it twenty-second, and reciprocal rank fusion let the blind channel cost the answer nothing.
Reality
- Evidence64
- Adoption15
- Hype gap−10
- Incentives60
- Confidence58
A dev.to post argues that IBM's chunkless RAG only helps on documents whose structure a parser can actually recover, and that on scanned PDFs and one-table wikis the agent ends up walking a tree the parser invented.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+34
- Incentives44
- Confidence36
An Apache-2.0 memory layer for coding agents returns a candidate only when several signals agree, and it labels every answer STRONG, WEAK or MISS so the calling agent has to branch on confidence before reusing an old fix.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence52
The generated descriptions were accurate and specific, and indexing them beside the table names dropped BM25's IDF for the term contact to 0.15, about what the same index gives words the tokenizer never strips out.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+12
- Incentives32
- Confidence57
The same platform ships regex and BM25 ranking in front of its tool catalog, while its memory store offers only depth, limit, page, path_prefix and view, so ranking memories is work the developer supplies.
Reality
- Evidence66
- Adoption
- Insufficient
- Hype gap−5
- Incentives70
- Confidence62
pgvector ships hnsw.ef_search at 40, the size of the candidate list its graph walk keeps in flight, and a top-20 query with a tenant filter can come back with two rows and no error to explain it.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+12
- Incentives20
- Confidence58
Perplexity pairs about 190 million web pages with 69,721 agent-written queries, which is closer to production than most retrieval tests get, and keeps the corpus, queries and labels private so it stays the only party able to run it.
Reality
- Evidence50
- Adoption15
- Hype gap+15
- Incentives88
- Confidence55
An independent rebuild of the SQLite side says the gap is real, but the published tables measure two different queries on two different graphs, so the depth at which SQLite's plan collapses on your data is a number you have to take yourself.
Reality
- Evidence55
- Adoption10
- Hype gap+30
- Incentives60
- Confidence50
The FAISS-plus-BM25 retrieval in this writeup does address vocabulary mismatch, but the agent loop around it ran at over 213 seconds a step on CPU, and that figure decided the deployment, not the retrieval design.
Reality
- Evidence34
- Adoption14
- Hype gap+42
- Incentives38
- Confidence56
A dev.to postmortem argues most RAG failures happen in retrieval. The debugging order it recommends survives scrutiny. The headline percentage it leans on does not.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+38
- Incentives30
- Confidence48
A build report claims 87.3% root-cause accuracy across 2,400 incident scenarios and a 59% cut in mean diagnosis time. The miss rate and the denominators deserve as much attention as the headline.
Reality
- Evidence30
- Adoption15
- Hype gap+38
- Incentives62
- Confidence52
Earlier coverage
- MCP's ceiling is the token bill, not the catalog: 75,000 servers nobody can afford to configure
Build · August 21, 2026 · 1 publisher
- BM25 lost worst on the queries full of file paths. In 54% of them, the path is not in the answer
Build · August 19, 2026 · 1 publisher
- The model was never the bottleneck: inside a Kannada RAG rebuild that blames retrieval
Build · August 18, 2026 · 1 publisher
- Your RAG Cannot Find SKU-4471, And A Bigger Embedding Model Will Not Help
Build · August 16, 2026 · 1 publisher
- Semantic code search over a monorepo is now a plumbing job, and the plumbing is the hard part
Build · August 15, 2026 · 1 publisher