Skip to content

Build1 publisher2 min readPublished

Every QueryFusionRetriever fusion mode in llama_index writes into the retriever's cached scores

llama_index's QueryFusionRetriever writes into retriever-owned score wrappers in all four fusion modes, two more than the first patch fixed. Over a caching retriever, later queries read corrupted cached scores, and so far only an issue covers reciprocal_rerank.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Every QueryFusionRetriever fusion mode in llama_index writes into the retriever's cached scores
Generated illustration

What happened

  • On October 3, michaelswissa reported that QueryFusionRetriever corrupts rankings when a caching retriever hands back shared NodeWithScore wrappers.
  • In his case, relative-score fusion normalized a shared wrapper for one result set, the next set summed it again, and the relevant document sank to -3.0 and ranked last.
  • In a per-query cache, simple fusion's dedup wrote another query's 0.9 over the 0.4 cached for the original query, with no exception and no change to the returned ranking.
  • The reviewer filed the two modes the first patch missed as #23351, and #23352 fixes the simple half.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A reciprocal_rerank pipeline over a caching retriever serves every later query from overwritten scores, and the call that caused the damage ranks correctly, so checks on returned rankings will miss it.
  • constraint A clean run of the #23332 shared-wrapper repro does not clear simple fusion. Validating that mode needs a per-query cache whose scores vary by query, as BM25 scores do.
  • decision Regression tests for fusion code have to inspect the retriever's cache after each call, because the tests in #23352 found the returned outputs were never wrong.

Each fusion mode gets a Dict[Tuple[str, int], List[NodeWithScore]] holding the result lists the retriever produced [8]. Those wrappers belong to the retriever. If the retriever caches them, any score assignment inside fusion edits the cache [8]. The bug only shows up with a retriever that keeps its wrappers between calls.

Every mode writes into the wrappers it receives [4][8]. The difference between them is whether anyone notices. Relative-score fusion breaks the current ranking, so it fails the first test anyone runs [9]. reciprocal_rerank and simple return a correct answer and damage only what the cache hands back next [10][13].

reciprocal_rerank computes ranks first. Only after that does it write fused scores back into the retriever's wrappers [5][10]. In the reviewer's repro, cached scores of 0.9, 0.1 and 0.8 came back as roughly 0.0333 and 0.0164. Every later retrieval from that cache reads those values [10].

simple fusion dedups by node hash and writes the max score into the first-seen wrapper [6]. It passes the shared-wrapper repro from #23332 for the most boring reason available: max(0.9, 0.9) is 0.9 [11]. "The obvious test passes, the mode looks clean, and the bug ships," the reviewer wrote [14]. The write only changes a value when one node hash arrives in distinct wrappers with different scores [12]. A per-query cache produces exactly that when scoring depends on the query, as with BM25 or query-weighted retrievers [12].

The post does not say which llama_index release carries #23333 or #23352, or when the reciprocal_rerank half will get a PR [1].

The post offers two fix shapes and calls both correct [18]. The one it spells out copies at the dispatch sites, so a single invariant, "fusion works on copies", covers every mode [17]. Each wrapper is copied once per occurrence in each list and never deduped by identity, so one object in two lists becomes two copies [17]. Deduping by identity would give two lists the same copy and bring the aliasing back [17]. I think this is the right fix for a library. The fusion functions receive state they do not own. One copy at the boundary covers all four modes with a single change.

The strongest check in the write-up is an ownership test. It sets the returned wrapper's score to 999.0, then asserts the cached score is still 0.4 [16]. The reviewer estimates the full audit at about 30 minutes per suspect function: grep for writes into inputs, assert on the caller's state, then test ownership [19].

What to watch

  • A fix PR for the reciprocal_rerank half of #23351, and whether it copies wrappers at the dispatch sites or patches that one function.
  • The first llama_index release that carries both #23333 and #23352.
  • Whether other llama_index components that post-process NodeWithScore lists turn out to write into their inputs too, since the reviewer says the bug class generalizes far past this retriever.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories