Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Debug the index before you swap the model: the RAG checklist works, the 73% does not

A dev.to postmortem argues most RAG failures happen in retrieval. The debugging order it recommends survives scrutiny. The headline percentage it leans on does not.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Debug the index before you swap the model: the RAG checklist works, the 73% does not
Generated illustration

What happened

  • A practitioner post on dev.to recounts chasing a confidently wrong RAG answer with a bigger model, a tuned prompt and a bolded stay-in-context instruction, with no improvement.
  • The author's diagnosis was that retrieval handed over the wrong chunks and the model summarised them faithfully.
  • Its first instruction is structural: prove you can re-chunk and re-embed the whole corpus without taking live search offline.
  • Its main chunking target is the 1,000-character split with 100 overlap, which cuts sentences, tables and functions apart.

Why it matters

  • decision It reorders the first hour of an incident. Reading ten chunks cold and inspecting the index come before prompt surgery, because they cost an afternoon and no infrastructure.
  • cost A model swap made to fix a retrieval fault is a permanent increase in per-request price bought to solve nothing, and the bill keeps arriving after the bug is found.
  • constraint Teams whose indexing and query paths share one system cannot change chunking without downtime, so the layer that fails silently becomes the layer they never revisit.
  • exposure Anyone quoting the 73% in a budget request or a postmortem is leaning on a figure with no named study behind it, and will be asked for one.

Bad chunks do not page anyone. A fixed-size splitter that cut a table mid-row still returns results with plausible similarity scores, and the generator still writes fluent prose over them [7][8]. That is why the model takes the blame: it is the only component in the chain whose output a human reads and can judge wrong on sight. Everything upstream fails quietly, which is also why the author's first three moves (bigger model, tuned prompt, a bolded instruction to stay in context) changed nothing [1][2].

The 73% figure is carrying more weight than it can hold. Take it at face value and retrieval failures outrun generation failures about 2.7 to 1 [21], which is a strong claim about where the first hour of a postmortem should go. The post credits it to "industry analysis in 2026" and does not name the analysis or say what counted as a failure [18][16]. A number with no denominator and no failure taxonomy cannot distinguish "retrieval returned nothing relevant" from "retrieval returned the right section with the wrong half of it".

The same thinness runs through the two upgrades the post recommends. Semantic chunking is credited with lifting accuracy meaningfully over fixed-size on the same dataset in a published comparison [19], and prepending a heading or short summary before embedding is said to measurably improve recall [20]. Neither comes with a figure.

The ordering still holds, and it holds on repair cost rather than on any percentage. Pulling ten random chunks and reading them cold needs no infrastructure and no eval harness [12]. The pass condition is legible to anyone: can this chunk answer a question by itself, or does it only make sense beside its neighbour [10]. A model swap, by contrast, raises the price of every request you will ever serve and does nothing about an index that never contained the answer.

The decoupling instruction is the one item with hard arithmetic behind it. Indexing can take minutes per document [3]; the query path is meant to finish in under about three seconds end to end [4]. At the friendliest reading of "minutes", the offline work for a single document consumes twenty times the entire online budget [17]. Share one system between them and every re-chunk becomes downtime, so chunking becomes the layer nobody touches [5][6], and chunking is the layer most likely to be broken [7]. Naive RAG was a prototype, and this is the part of it that ossifies first [14].

The honest version of this checklist has no percentage in it at all. Cheap diagnostics go first because they are cheap, and the cheap ones happen to sit in retrieval. The 73% is decoration on a decision your own repair bill already makes.

What to watch

  • A named study with a stated failure taxonomy behind the 73% would turn a repeated number into something a planner can use.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence26
Adoption
Insufficient
Hype gap+38
Incentives30
Confidence48
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The author's first confidently wrong RAG answer led him to swap in a bigger model, tune the prompt, and add "only answer from the context provided" in bold; the answer got no better.

    ReportedSupportedView cited source
  2. [2]

    The author concluded the model was faithfully summarizing the context it was handed; retrieval had returned the wrong chunks, so the system answered the wrong question fluently.

    ReportedSupportedView cited source
  3. [3]

    The offline indexing path parses the source, cleans the text, chunks it, optionally enriches each chunk with context, embeds it, and writes to a vector store and a keyword index; it can take minutes per document and runs in the background.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 24, 2026

    The Retrieval Checklist I Wish I'd Had Before Shipping RAG

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories