One developer's test on 484 SEPA rulebook passages found that plain-English questions push several answers out of a top-5 vector search. Because the test measures each answer's rank directly, the failure shows up in retrieval, before the language model writes anything.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence45
Amazon OpenSearch Service needs about 1.3 TB of resident RAM for 100 million 1,536-dimension FP32 vectors with one replica, a dev.to sizing post calculates. Raw vector values are about 98% of each entry, so the encoding picked before ingestion decides most of that memory and the node count behind it.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence62
Databricks made Lakebase Search generally available, pairing BM25 with a Postgres vector index it says runs 4 times cheaper than pgvector at 100M vectors. The vector index lives in object storage behind a cache, so large agent corpora no longer need RAM sized to the whole index.
Publishers:databricks.com · neon.tech Reality
- Evidence40
- Adoption25
- Hype gap+35
- Incentives90
- Confidence50
Cloudflare made AI Search generally available with native image embeddings and PDF OCR, and will start billing for it on November 1, 2026. Which embedding model an instance runs now decides whether an image query searches pixels or a caption.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives85
- Confidence45
Cohere's Embed 5 lets teams index with Pro at $0.12 per million text tokens and query with Fast at $0.08 against the same vectors. The quality and throughput figures behind that split come from Cohere's own tests, run on datasets and parsing pipelines a buyer may not share.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives75
- Confidence55
AWS says metadata pre-filtering in S3 Vectors returns up to 5x more matching vectors on highly selective filters, at no additional cost. Existing indexes keep CLASSIC search until updated, so multi-tenant RAG stores get the gain only after an index-mode switch.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives80
- Confidence60
AWS added a SearchVectors API that keeps embeddings in the same table as the data they describe. The second store and the job that feeds it are now optional, and metered three ways.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence62
Throughline's on-call agent returns a three-way coverage verdict with every memory recall, so a timed-out search cannot pass as "no prior incidents". Its author hit the same error-as-empty trap in CockroachDB's managed MCP server, where failures come back as HTTP 200.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap−10
- Incentives30
- Confidence40
Exact-substring scoring put a developer's 700-line RAG tool at a 65% retrieval hit-rate at k=3, 13 points below what a token-overlap scorer found. The bug also made extra retrieved chunks look worthless, and a 20-question test set added 15 points of noise.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence55
Sentinel's builder trained its alert model on fraud versus cleared cases, reaching 0.9465 AUC after three passes that over-called fraud on every case. The finished agent lets an LLM write the explanation and leaves every action to fixed bank rules.
Reality
- Evidence35
- Adoption5
- Hype gap+10
- Incentives55
- Confidence40
A dev.to walkthrough of a permission-aware Postgres project puts the ACL test in a CTE that the vector ranking reads from, so the nearest-neighbour search only ever orders rows the caller may see.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+10
- Incentives35
- Confidence66
A dev.to walkthrough replaces a managed vector store and a hosted embedding API with sqlite-vec and a local model on port 11434. Its schema declares float[768], so changing model dimension means re-embedding everything.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+25
- Incentives22
- Confidence48
A dev.to writeup by nasiko_labs describes a semantic memory tier whose first version inserted a row per session, so one agent could read "cautious" and "aggressive" at once. The rewrite keys each belief by fact type and archives the superseded version.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence45
A dev.to post lays out a four-layer recommender and a 100 millisecond p99 budget with no slot for a generative call in the hot path. The figures are its author's own allocation, not a measurement from a running system.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence45
A dev.to test stores each shop's position as a unit-sphere triple and gets ordering that matches haversine to about half a metre across seven places, while the geohash index it was compared against missed the nearest shop.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−10
- Incentives45
- Confidence55
Shrijith Venkatramana's embedding evaluation guide treats retrieval as a ranking problem scored with Recall@k, MRR and NDCG. Every figure in it is a worked example, and the labelling is the real bill.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+8
- Incentives50
- Confidence52
A new AWS walkthrough moves a three-backend healthcare agent off ECS with Fargate by wrapping the same code in a runtime decorator. It changes the reasoning model in the same step, so the two versions cannot be diffed for cost.
Reality
- Evidence42
- Adoption15
- Hype gap+35
- Incentives80
- Confidence55
A dev.to post by Rijul argues that semantic similarity is the wrong tool for error codes, part numbers and filenames, and sketches a retrieval pipeline that runs keyword and vector search side by side. It reports no measurements.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence58
The first VDPU samples are back from fab and go into system evaluation in the fourth quarter. Everything Dnotitia has published so far, including the 5.77x throughput claim, was measured on a four-card FPGA rig.
Reality
- Evidence38
- Adoption12
- Hype gap+22
- Incentives78
- Confidence52
The paper reports 2 to 1,000 times the throughput of prior filtered-search methods at fixed recall, and says an existing HNSW library is enough to implement it. Both claims turn on how the graph is built.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+35
- Incentives65
- Confidence50
Earlier coverage
- JSONB and pgvector cover two of the five roles this Postgres consolidation absorbs
Build · September 17, 2026 · 1 publisher
- Co-locating embeddings with permissions collapses the RAG fetch into one SQL statement
Build · September 17, 2026 · 1 publisher
- Hashing chunk IDs into five shard keys widens a DynamoDB vector search to 500 candidates
Build · September 16, 2026 · 1 publisher
- One Lambda fans a query into a prefix match and a 512-dimension cosine search
Build · September 15, 2026 · 1 publisher
- Valkey 9.1 cuts per-key string overhead by 17 to 44 percent with no config change
Build · September 12, 2026 · 1 publisher
- A boolean stream flag leaves callers guessing whether they get a string or a generator
Build · September 12, 2026 · 1 publisher
- The server injects up to five project notes before the agent takes its first turn
Build · September 11, 2026 · 1 publisher
- Cosine similarity ranks the menu above the peanut allergy note
Build · September 11, 2026 · 1 publisher
- A Postgres default of 40 caps how many neighbors your vector search can return
Build · September 10, 2026 · 1 publisher
- Adjudicating every extracted fact against its five nearest memories costs one LLM call apiece
Build · September 6, 2026 · 1 publisher
- Only the corpus shrinks a billion-vector face index without costing recall
Build · August 27, 2026 · 1 publisher
- Four control planes, one Postgres: a team's case against polyglot persistence
Build · August 21, 2026 · 1 publisher
- DuckDB's vss extension removes a database from your RAG stack, then names the price
Build · August 15, 2026 · 1 publisher