Barclays expects half its developers to use Anthropic's Claude Code by the end of 2026 and most of them during 2027. Its published results so far come from a staff search tool and an email sorter, so the coding goal is still a seat count.
Reality
- Evidence40
- Adoption45
- Hype gap+30
- Incentives80
- Confidence60
Plain RAG and GraphRAG got none of 31 counting and superlative questions right on a 100-question TigerGraph hackathon benchmark, an entrant reports. A COUNT from the graph, set beside the evidence actually read, shows when an answer is incomplete.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence50
llama_index's QueryFusionRetriever writes into retriever-owned score wrappers in all four fusion modes, two more than the first patch fixed. Over a caching retriever, later queries read corrupted cached scores, and so far only an issue covers reciprocal_rerank.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives25
- Confidence55
Ai2 open-sourced AstaBrief 8B, which writes cited research reports in 51.1 seconds against 178.5 for Asta's Claude-powered mode. Labs that cannot send unpublished research questions to a hosted model can now run a cited-report generator on their own servers.
Perspective Coverage
3 publishers
- Builder
- Builder 55%
- Operator
- Operator 33%
- Investor
- Investor 12%
Reality
- Evidence55
- Adoption18
- Hype gap+22
- Incentives55
- Confidence60
One developer's test on 484 SEPA rulebook passages found that plain-English questions push several answers out of a top-5 vector search. Because the test measures each answer's rank directly, the failure shows up in retrieval, before the language model writes anything.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence45
OpenAI retired the Assistants API on August 26, 2026, without an automatic tool to move old Threads into the Conversation objects that replace them. Teams that relied on managed threads now write that migration and set how much history each turn sends.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
A developer's GraphRAG build turns Thailand's 96-section PDPA into a 190-node graph so an agent can walk from Section 26 to the fines that cite it. The published part argues the vector-search gap from vocabulary and reports no measured recall figures.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives25
- Confidence45
Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
Notion bought ZeroEntropy after its reranker cut latency in Notion's reranking step by 85%, by Notion's count. The hosted product is being sunset with an open-weight release, so teams that built on ZeroEntropy's API now have to decide where the model runs.
Reality
- Evidence40
- Adoption40
- Hype gap+15
- Incentives60
- Confidence45
Cloudflare made AI Search generally available with native image embeddings and PDF OCR, and will start billing for it on November 1, 2026. Which embedding model an instance runs now decides whether an image query searches pixels or a caption.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives85
- Confidence45
GroundedDocs, a FastAPI RAG service, verifies Keycloak-signed tokens that expire in about 15 minutes and holds no key that can sign one. Keycloak now keeps the passwords and the signing key, so the identity provider becomes the system worth attacking.
Reality
- Evidence55
- Adoption8
- Hype gap+5
- Incentives20
- Confidence45
Cohere's Embed 5 lets teams index with Pro at $0.12 per million text tokens and query with Fast at $0.08 against the same vectors. The quality and throughput figures behind that split come from Cohere's own tests, run on datasets and parsing pipelines a buyer may not share.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives75
- Confidence55
AWS says metadata pre-filtering in S3 Vectors returns up to 5x more matching vectors on highly selective filters, at no additional cost. Existing indexes keep CLASSIC search until updated, so multi-tenant RAG stores get the gain only after an index-mode switch.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives80
- Confidence60
KAIST and Microsoft researchers say a map of named entities cut a search agent's input tokens from 206,500 to 88,100 per query on EnterpriseRAG-Bench. Whether that saving holds outside the benchmark depends on what the map costs to build and keep current.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence45
Cache-augmented generation costs about what retrieval does when the corpus is roughly 10 times the tokens retrieval would send, a dev.to analysis finds. Sparse traffic breaks the rule, because each query then pays the cache-write premium and caching becomes the most expensive option.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
A TigerGraph hackathon entry raised exact match from 67% to 99% on 100 questions, with the agent alone accounting for 3 of the 32 points. The rest needed a parser that made Wikipedia infobox fields countable in the graph.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence45
Microsoft's MarkItDown returns a single newline and exit code 0 for a scanned PDF, according to a dev.to test of version 0.1.8 on 14 files. Any pipeline that feeds a retrieval index from it has to check the returned text itself before writing a file.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence55
Memory handoffs cut replaced agent context by 98.65% across 10,241 completed runs, according to an audit published on dev.to. For agent builders, the open risk now sits in whether the agent recalls the right stored text, since lossy summaries are no longer the failure point.
Reality
- Evidence55
- Adoption15
- Hype gap+10
- Incentives
- Insufficient
- Confidence45
AWS added a SearchVectors API that keeps embeddings in the same table as the data they describe. The second store and the job that feeds it are now optional, and metered three ways.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence62
LiveReview's Maneshwar says Gemini File Search cost over 50 cents a review run because a reasoning model performed each search. A local index and a cheaper model brought runs down to 4 cents.
Reality
- Evidence45
- Adoption10
- Hype gap+15
- Incentives35
- Confidence40
Earlier coverage
- Checking a revision register keeps superseded documents out of RAG answers
Build · September 26, 2026 · 1 publisher
- Frontier models score at most 0.17 on a source-trust test that a two-line rule passes perfectly
Build · September 26, 2026 · 1 publisher
- Per-byte metering on 50KB records pushed Perplexity from DynamoDB to a home-built Rust store
Build · September 26, 2026 · 1 publisher
- A substring-match scorer made a weekend RAG build look 13 points worse than it was
Build · September 25, 2026 · 1 publisher
- Bedrock's managed video search embeds your footage in 512 dimensions every four seconds
Build · September 10, 2026 · 1 publisher
- Serving RAG retrieval as a cacheable GET lets an edge proxy absorb repeat lookups
Build · September 25, 2026 · 1 publisher
- Dropping the vector database hands chunk boundaries to Google's whitespace rule
Build · September 22, 2026 · 1 publisher
- Choosing the embedding model first locks the vec0 table to a fixed 768-dimension schema
Build · September 22, 2026 · 1 publisher
- A dedicated vector database adds a sync job to every write
Build · September 22, 2026 · 1 publisher
- Adding both the guard and sandbox moved measured attack success from 16% to 20%
Build · September 21, 2026 · 1 publisher
- JudgeStack stores today's ban and the dated announcement that imposed it as separate documents
Build · September 21, 2026 · 1 publisher
- A 100K-token agent session prefills 3M tokens over thirty turns without prefix reuse
Build · September 21, 2026 · 1 publisher
- A fifteen-line abstention rule removed more correct answers than confident wrong ones
Build · September 21, 2026 · 1 publisher
- Pen Test Partners' AI toaster broke its own CTF rules until the password moved into code
Security · September 20, 2026 · 1 publisher
- Firing 4.8% of the weights per token still leaves 125GB to keep resident
Build · September 19, 2026 · 1 publisher
- Storing a license ruling as a directional from-into pair stops a model inverting the verdict
Build · September 19, 2026 · 1 publisher
- Grading a retriever starts with hand-labelling 500 queries against 100,000 chunks
Build · September 19, 2026 · 1 publisher
- Stamping a trap chunk id into every negative test turned two false passes red
Build · September 19, 2026 · 1 publisher
- GitHub's review rule for AI code stops where you can explain and own the outcome
Build · September 18, 2026 · 1 publisher
- One error code explains why RAG pipelines keep a keyword index
Build · September 18, 2026 · 1 publisher
- Dnotitia's retrieval ASIC needs 1.73x its FPGA prototype to reach a 10x target
Build · September 18, 2026 · 1 publisher
- Aurora pgvector caps Bedrock Knowledge Bases at 2,000 dimensions in single precision
Build · September 17, 2026 · 1 publisher
- pgvector's cost-based planner picks approximate scans where exact scans hit perfect recall
Build · September 17, 2026 · 1 publisher
- Reading faithfulness against context recall tells you which half of a RAG pipeline broke
Build · September 17, 2026 · 1 publisher
- Co-locating embeddings with permissions collapses the RAG fetch into one SQL statement
Build · September 17, 2026 · 1 publisher
- Persistent memory and MCP tools make 27B enough for a local assistant on 24 GB
Build · September 17, 2026 · 1 publisher
- Hashing chunk IDs into five shard keys widens a DynamoDB vector search to 500 candidates
Build · September 16, 2026 · 1 publisher
- Chunkless RAG relocates the retrieval decision into the layout parser
Build · September 16, 2026 · 1 publisher
- 88 KB of read-only rows priced Aurora Serverless v2 out of an agentic RAG rewrite
Build · September 16, 2026 · 1 publisher
- A typed internal DSL trades first-try compile rate for fewer invented keywords
Build · September 16, 2026 · 1 publisher
- Every unanswerable question cleared the 0.35 refusal threshold by at least 0.09
Build · September 15, 2026 · 1 publisher
- Fine-tuning requests usually mean one of two things: missing knowledge or wrong style
Build · September 15, 2026 · 1 publisher
- A RAG design re-reads the user's department from Postgres before every vector search
Build · September 14, 2026 · 1 publisher
- Three judges score every answer in a RAG sweep, including one from the generator's own vendor
Build · September 14, 2026 · 1 publisher
- AiSAQ holds query-time RAM at 10 MB on a billion-vector index by moving PQ codes to SSD
Build · September 14, 2026 · 1 publisher
- Red Hat's AutoRAG preview turns chunking and embedding choices into a scored search
Product · September 14, 2026 · 1 publisher
- Toast 1 claims MTEB parity with OpenAI. The cost math in the pitch is off by 1000x.
Build · August 14, 2026 · 1 publisher
- "Infinite context" is not a spec: a four-task harness for testing agent memory
Build · August 16, 2026 · 1 publisher
- Numbering six retrieved chunks turned one handbook into three agreeing sources
Build · September 13, 2026 · 1 publisher
- A support agent's faithfulness check passed on documents from 2024 and 2023
Build · September 13, 2026 · 1 publisher