BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Debug the index before you swap the model: the RAG checklist works, the 73% does not
A dev.to postmortem argues most RAG failures happen in retrieval. The debugging order it recommends survives scrutiny. The headline percentage it leans on does not.
The Engineer · Build desk

What happened
- A practitioner post on dev.to recounts chasing a confidently wrong RAG answer with a bigger model, a tuned prompt and a bolded stay-in-context instruction, with no improvement.
- The author's diagnosis was that retrieval handed over the wrong chunks and the model summarised them faithfully.
- Its first instruction is structural: prove you can re-chunk and re-embed the whole corpus without taking live search offline.
- Its main chunking target is the 1,000-character split with 100 overlap, which cuts sentences, tables and functions apart.
Why it matters
- decision It reorders the first hour of an incident. Reading ten chunks cold and inspecting the index come before prompt surgery, because they cost an afternoon and no infrastructure.
- cost A model swap made to fix a retrieval fault is a permanent increase in per-request price bought to solve nothing, and the bill keeps arriving after the bug is found.
- constraint Teams whose indexing and query paths share one system cannot change chunking without downtime, so the layer that fails silently becomes the layer they never revisit.
- exposure Anyone quoting the 73% in a budget request or a postmortem is leaning on a figure with no named study behind it, and will be asked for one.
Bad chunks do not page anyone. A fixed-size splitter that cut a table mid-row still returns results with plausible similarity scores, and the generator still writes fluent prose over them [7][8]. That is why the model takes the blame: it is the only component in the chain whose output a human reads and can judge wrong on sight. Everything upstream fails quietly, which is also why the author's first three moves (bigger model, tuned prompt, a bolded instruction to stay in context) changed nothing [1][2].
The 73% figure is carrying more weight than it can hold. Take it at face value and retrieval failures outrun generation failures about 2.7 to 1 [21], which is a strong claim about where the first hour of a postmortem should go. The post credits it to "industry analysis in 2026" and does not name the analysis or say what counted as a failure [18][16]. A number with no denominator and no failure taxonomy cannot distinguish "retrieval returned nothing relevant" from "retrieval returned the right section with the wrong half of it".
The same thinness runs through the two upgrades the post recommends. Semantic chunking is credited with lifting accuracy meaningfully over fixed-size on the same dataset in a published comparison [19], and prepending a heading or short summary before embedding is said to measurably improve recall [20]. Neither comes with a figure.
The ordering still holds, and it holds on repair cost rather than on any percentage. Pulling ten random chunks and reading them cold needs no infrastructure and no eval harness [12]. The pass condition is legible to anyone: can this chunk answer a question by itself, or does it only make sense beside its neighbour [10]. A model swap, by contrast, raises the price of every request you will ever serve and does nothing about an index that never contained the answer.
The decoupling instruction is the one item with hard arithmetic behind it. Indexing can take minutes per document [3]; the query path is meant to finish in under about three seconds end to end [4]. At the friendliest reading of "minutes", the offline work for a single document consumes twenty times the entire online budget [17]. Share one system between them and every re-chunk becomes downtime, so chunking becomes the layer nobody touches [5][6], and chunking is the layer most likely to be broken [7]. Naive RAG was a prototype, and this is the part of it that ossifies first [14].
The honest version of this checklist has no percentage in it at all. Cheap diagnostics go first because they are cheap, and the cheap ones happen to sit in retrieval. The 73% is decoration on a decision your own repair bill already makes.
What to watch
- A named study with a stated failure taxonomy behind the 73% would turn a repeated number into something a planner can use.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+38
- Incentives30
- Confidence48
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The author's first confidently wrong RAG answer led him to swap in a bigger model, tune the prompt, and add "only answer from the context provided" in bold; the answer got no better.
- [2]
The author concluded the model was faithfully summarizing the context it was handed; retrieval had returned the wrong chunks, so the system answered the wrong question fluently.
- [3]
The offline indexing path parses the source, cleans the text, chunks it, optionally enriches each chunk with context, embeds it, and writes to a vector store and a keyword index; it can take minutes per document and runs in the background.
- [4]
The online query path runs on every user request under a latency budget the post puts at under roughly 3 seconds end to end, covering optional query rewriting, candidate retrieval, reranking, prompt assembly with citations, generation, and trace logging.
- [5]
The post calls coupling the indexing and query paths the most common architectural mistake: if re-indexing forces the query path offline, you cannot iterate on chunking or swap embedding models without downtime, so you stop iterating.
- [6]
The post's first check is whether you can re-chunk and re-embed the whole corpus without taking live search down, and it says to decouple the paths before anything else if you cannot.
- [7]
The post says chunking is where pipelines silently fail, because bad chunks do not throw errors; they return technically-relevant, practically-useless context.
- [8]
The naive chunking default described is splitting every 1,000 characters with 100 overlap, which cuts sentences mid-thought, tables mid-row, and code mid-function.
- [9]
Structure-aware splitting uses the document's own boundaries: ## headings for docs, per-function or per-class for code, per-row for tables; the post rates it low effort with a big payoff.
- [10]
The post's rule is that each chunk should be able to answer a question on its own; if a chunk only makes sense next to its neighbour, the splitting is too aggressive.
- [11]
Chunks that are too small fragment ideas and chunks that are too large dilute the signal, forcing the model to average across mostly-irrelevant text.
- [12]
The post's chunking check is to pull ten random chunks and read them cold, asking whether each stands on its own or is a sentence fragment or orphaned table row.
- [13]
The post recommends storing arriving metadata alongside the chunk: author, date, source, section, product version, document type, access level.
- [14]
The post describes naive RAG ("chunk, embed, cosine similarity, stuff into prompt") as always having been a prototype.
- [15]
The material is a single-author practitioner post published on dev.to under the headline "The Retrieval Checklist I Wish I'd Had Before Shipping RAG".
- [16]
The source attaches no named study, dataset, or definition of "failure" to the 73% figure: the count of named citations supporting it is zero.
- [17]
Reading "minutes per document" at its lowest value, one minute, the offline indexing of a single document consumes about 20 times the entire end-to-end budget for a live query.
- [18]
The post states that industry analysis in 2026 keeps landing on the same number: when RAG fails, the failure is in retrieval roughly 73% of the time, not generation.
ReportedInsufficientSource: dev.to post by James Anderson, crediting unnamed "industry analysis in 2026"View cited source - [19]
Semantic chunking computes sentence-to-sentence similarity and starts a new chunk where meaning shifts; the post says it costs more compute and that a published comparison reported it lifting accuracy meaningfully over fixed-size on the same dataset.
- [20]
The post says embedding only raw body text strips situating context, and recommends prepending the heading, a short document summary, or a one-line description before embedding, the core idea behind "contextual retrieval", which it says measurably improves recall.
- [21]
If retrieval accounts for 73% of RAG failures, retrieval failures outnumber generation failures by about 2.7 to 1.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toThe Retrieval Checklist I Wish I'd Had Before Shipping RAG
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.