Skip to content

Topic

Retrieval-Augmented Generation

Ingestion, chunking, retrieval, and grounding pipelines that feed retrieved document context to an LLM.

Current stories

build3 publishers

Ai2's open AstaBrief 8B writes cited research reports 3.5 times faster than Asta's Claude mode

Ai2 open-sourced AstaBrief 8B, which writes cited research reports in 51.1 seconds against 178.5 for Asta's Claude-powered mode. Labs that cannot send unpublished research questions to a hosted model can now run a cited-report generator on their own servers.

Perspective Coverage

3 publishers
Builder
Builder 55%
Operator
Operator 33%
Investor
Investor 12%

Reality

Evidence55
Adoption18
Hype gap+22
Incentives55
Confidence60
build1 publisher

Plain-language questions push SEPA rulebook answers out of a top-5 vector search

One developer's test on 484 SEPA rulebook passages found that plain-English questions push several answers out of a top-5 vector search. Because the test measures each answer's rank directly, the failure shows up in retrieval, before the language model writes anything.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap0
Incentives
Insufficient
Confidence45
build1 publisher

Crutches built from measured failures lift a local Qwen 3B from 33% to 52% on post-cutoff facts

Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence40
build1 publisher

Cache expiry decides whether a whole knowledge base in the prompt undercuts retrieval

Cache-augmented generation costs about what retrieval does when the corpus is roughly 10 times the tokens retrieval would send, a dev.to analysis finds. Sparse traffic breaks the rule, because each query then pays the cache-write premium and caching becomes the most expensive option.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence40
build1 publisher

MarkItDown reports success on scanned PDFs that yield only a newline

Microsoft's MarkItDown returns a single newline and exit code 0 for a scanned PDF, according to a dev.to test of version 0.1.8 on 14 files. Any pipeline that feeds a retrieval index from it has to check the returned text itself before writing a file.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap0
Incentives
Insufficient
Confidence55

Earlier coverage

  1. Checking a revision register keeps superseded documents out of RAG answers

    Build · September 26, 2026 · 1 publisher

  2. Frontier models score at most 0.17 on a source-trust test that a two-line rule passes perfectly

    Build · September 26, 2026 · 1 publisher

  3. Per-byte metering on 50KB records pushed Perplexity from DynamoDB to a home-built Rust store

    Build · September 26, 2026 · 1 publisher

  4. A substring-match scorer made a weekend RAG build look 13 points worse than it was

    Build · September 25, 2026 · 1 publisher

  5. Bedrock's managed video search embeds your footage in 512 dimensions every four seconds

    Build · September 10, 2026 · 1 publisher

  6. Serving RAG retrieval as a cacheable GET lets an edge proxy absorb repeat lookups

    Build · September 25, 2026 · 1 publisher

  7. Dropping the vector database hands chunk boundaries to Google's whitespace rule

    Build · September 22, 2026 · 1 publisher

  8. Choosing the embedding model first locks the vec0 table to a fixed 768-dimension schema

    Build · September 22, 2026 · 1 publisher

  9. A dedicated vector database adds a sync job to every write

    Build · September 22, 2026 · 1 publisher

  10. Adding both the guard and sandbox moved measured attack success from 16% to 20%

    Build · September 21, 2026 · 1 publisher

  11. JudgeStack stores today's ban and the dated announcement that imposed it as separate documents

    Build · September 21, 2026 · 1 publisher

  12. A 100K-token agent session prefills 3M tokens over thirty turns without prefix reuse

    Build · September 21, 2026 · 1 publisher

  13. A fifteen-line abstention rule removed more correct answers than confident wrong ones

    Build · September 21, 2026 · 1 publisher

  14. Pen Test Partners' AI toaster broke its own CTF rules until the password moved into code

    Security · September 20, 2026 · 1 publisher

  15. Firing 4.8% of the weights per token still leaves 125GB to keep resident

    Build · September 19, 2026 · 1 publisher

  16. Storing a license ruling as a directional from-into pair stops a model inverting the verdict

    Build · September 19, 2026 · 1 publisher

  17. Grading a retriever starts with hand-labelling 500 queries against 100,000 chunks

    Build · September 19, 2026 · 1 publisher

  18. Stamping a trap chunk id into every negative test turned two false passes red

    Build · September 19, 2026 · 1 publisher

  19. GitHub's review rule for AI code stops where you can explain and own the outcome

    Build · September 18, 2026 · 1 publisher

  20. One error code explains why RAG pipelines keep a keyword index

    Build · September 18, 2026 · 1 publisher

  21. Dnotitia's retrieval ASIC needs 1.73x its FPGA prototype to reach a 10x target

    Build · September 18, 2026 · 1 publisher

  22. Aurora pgvector caps Bedrock Knowledge Bases at 2,000 dimensions in single precision

    Build · September 17, 2026 · 1 publisher

  23. pgvector's cost-based planner picks approximate scans where exact scans hit perfect recall

    Build · September 17, 2026 · 1 publisher

  24. Reading faithfulness against context recall tells you which half of a RAG pipeline broke

    Build · September 17, 2026 · 1 publisher

  25. Co-locating embeddings with permissions collapses the RAG fetch into one SQL statement

    Build · September 17, 2026 · 1 publisher

  26. Persistent memory and MCP tools make 27B enough for a local assistant on 24 GB

    Build · September 17, 2026 · 1 publisher

  27. Hashing chunk IDs into five shard keys widens a DynamoDB vector search to 500 candidates

    Build · September 16, 2026 · 1 publisher

  28. Chunkless RAG relocates the retrieval decision into the layout parser

    Build · September 16, 2026 · 1 publisher

  29. 88 KB of read-only rows priced Aurora Serverless v2 out of an agentic RAG rewrite

    Build · September 16, 2026 · 1 publisher

  30. A typed internal DSL trades first-try compile rate for fewer invented keywords

    Build · September 16, 2026 · 1 publisher

  31. Every unanswerable question cleared the 0.35 refusal threshold by at least 0.09

    Build · September 15, 2026 · 1 publisher

  32. Fine-tuning requests usually mean one of two things: missing knowledge or wrong style

    Build · September 15, 2026 · 1 publisher

  33. A RAG design re-reads the user's department from Postgres before every vector search

    Build · September 14, 2026 · 1 publisher

  34. Three judges score every answer in a RAG sweep, including one from the generator's own vendor

    Build · September 14, 2026 · 1 publisher

  35. AiSAQ holds query-time RAM at 10 MB on a billion-vector index by moving PQ codes to SSD

    Build · September 14, 2026 · 1 publisher

  36. Red Hat's AutoRAG preview turns chunking and embedding choices into a scored search

    Product · September 14, 2026 · 1 publisher

  37. Toast 1 claims MTEB parity with OpenAI. The cost math in the pitch is off by 1000x.

    Build · August 14, 2026 · 1 publisher

  38. "Infinite context" is not a spec: a four-task harness for testing agent memory

    Build · August 16, 2026 · 1 publisher

  39. Numbering six retrieved chunks turned one handbook into three agreeing sources

    Build · September 13, 2026 · 1 publisher

  40. A support agent's faithfulness check passed on documents from 2024 and 2023

    Build · September 13, 2026 · 1 publisher