Skip to content

Build1 publisher3 min readPublished

Your retrieval stack knows who owns retry logic. It cannot tell you what keeps breaking.

Lookup and aggregation are opposite retrieval problems, and chunk similarity only answers one of them. The second failure is architectural, not a tuning bug.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A retrieval system pointed at five years of a team's engineering documents, including design docs, incident postmortems and architecture decision records, gives a pretty accurate and well-cited answer to "which service owns the payments retry logic".
  • Asked which failure causes recur most often across all the postmortems, answer quality goes down; depending on the setup the system may list a handful of incidents that happen to use the word "recurring", giving no idea of the underlying pattern.
  • Both questions can look similar from the outside, but architecturally they are opposites.
  • The first question's answer can be found in a specific document, which is precisely what similarity search was built for.
  • The second question's answer shows up only after the entire collection has been surveyed and understood, which requires a completely different retrieval mechanism.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A ByteByteGo explainer on GraphRAG opens with a scenario worth stealing for your next design review: a retrieval system pointed at five years of design docs, incident postmortems and architecture decision records answers "which service owns the payments retry logic" accurately and with citations, then degrades badly on "which failure causes recur most often across all the postmortems" [s1c1][s1c2]. The second failure is not a tuning problem, which matters because teams routinely spend a quarter swapping embedding models and adding rerankers without moving that answer at all.

The two questions look alike from the outside. Architecturally they are opposites [s1c3]. The first has an answer sitting inside one specific document, which is precisely what similarity search was built for [s1c4]. The second has an answer that only exists after the entire collection has been surveyed, and that requires a different retrieval mechanism [s1c5].

The mechanism explains the split. Standard RAG slices documents into chunks of a few hundred to a few thousand tokens, embeds each chunk as a vector, and stores those vectors in an index; at query time the question becomes a vector, the index returns the nearest chunk vectors, and their original text goes into the prompt [s1c6][s1c7]. The whole design rests on one assumption: that text answering a question resembles the question [s1c8]. For the ownership query that assumption holds, because the question shares vocabulary with the architecture decision record where ownership was recorded [s1c9].

Microsoft's GraphRAG documentation draws the same line with different words, separating local queries, whose answers resemble the query and live in a small number of text regions, from global queries, which require reasoning across large portions of a dataset or all of it [s1c10][s1c11]. Ownership is local. Recurrence across all postmortems is global [s1c12].

What happens on the global query is worth being precise about. "Recur most often" produces a vector, and the index returns whatever is nearest, which across a corpus of incident reports means documents using words like "recurring" or "frequent" [s1c13]. That is a coincidence of vocabulary, not a count [s1c13]. Nothing in the pipeline represents how many documents share a root cause, and a nearest-neighbour lookup returns roughly the same handful of chunks whether three postmortems or ninety mention the same failure [s1d1]. The observed output, a list of incidents that happen to contain the word "recurring" with no underlying pattern, is exactly what the architecture should be expected to produce [s1c2].

GraphRAG is aimed at the second class of question [s1c14]. The same explainer walks through knowledge graph construction from ordinary documents, an indexing pipeline, community detection with hierarchical summaries, and separate local and global search paths [s1c15]. It also promises sections on cost, latency and maintenance tradeoffs, and on when standard RAG remains the better option [s1c16]. That last section is the honest signal. The indexing work is paid across the whole corpus, while the benefit lands only on the global slice of your query traffic [s1d2].

Before anyone proposes a graph pipeline, classify a month of real queries into local and global. The author notes the post is assembled from publicly shared details rather than first-hand benchmarks [s1c17], so the numbers that decide this are the ones from your own logs: the global share of traffic, and what re-indexing five years of documents costs every time they change.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories