Build1 publisher2 min readPublished
Every QueryFusionRetriever fusion mode in llama_index writes into the retriever's cached scores
llama_index's QueryFusionRetriever writes into retriever-owned score wrappers in all four fusion modes, two more than the first patch fixed. Over a caching retriever, later queries read corrupted cached scores, and so far only an issue covers reciprocal_rerank.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- On October 3, michaelswissa reported that QueryFusionRetriever corrupts rankings when a caching retriever hands back shared NodeWithScore wrappers.
- In his case, relative-score fusion normalized a shared wrapper for one result set, the next set summed it again, and the relevant document sank to -3.0 and ranked last.
- In a per-query cache, simple fusion's dedup wrote another query's 0.9 over the 0.4 cached for the original query, with no exception and no change to the returned ranking.
- The reviewer filed the two modes the first patch missed as #23351, and #23352 fixes the simple half.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure A reciprocal_rerank pipeline over a caching retriever serves every later query from overwritten scores, and the call that caused the damage ranks correctly, so checks on returned rankings will miss it.
- constraint A clean run of the #23332 shared-wrapper repro does not clear simple fusion. Validating that mode needs a per-query cache whose scores vary by query, as BM25 scores do.
- decision Regression tests for fusion code have to inspect the retriever's cache after each call, because the tests in #23352 found the returned outputs were never wrong.
Each fusion mode gets a Dict[Tuple[str, int], List[NodeWithScore]] holding the result lists the retriever produced [8]. Those wrappers belong to the retriever. If the retriever caches them, any score assignment inside fusion edits the cache [8]. The bug only shows up with a retriever that keeps its wrappers between calls.
Every mode writes into the wrappers it receives [4][8]. The difference between them is whether anyone notices. Relative-score fusion breaks the current ranking, so it fails the first test anyone runs [9]. reciprocal_rerank and simple return a correct answer and damage only what the cache hands back next [10][13].
reciprocal_rerank computes ranks first. Only after that does it write fused scores back into the retriever's wrappers [5][10]. In the reviewer's repro, cached scores of 0.9, 0.1 and 0.8 came back as roughly 0.0333 and 0.0164. Every later retrieval from that cache reads those values [10].
simple fusion dedups by node hash and writes the max score into the first-seen wrapper [6]. It passes the shared-wrapper repro from #23332 for the most boring reason available: max(0.9, 0.9) is 0.9 [11]. "The obvious test passes, the mode looks clean, and the bug ships," the reviewer wrote [14]. The write only changes a value when one node hash arrives in distinct wrappers with different scores [12]. A per-query cache produces exactly that when scoring depends on the query, as with BM25 or query-weighted retrievers [12].
The post does not say which llama_index release carries #23333 or #23352, or when the reciprocal_rerank half will get a PR [1].
The post offers two fix shapes and calls both correct [18]. The one it spells out copies at the dispatch sites, so a single invariant, "fusion works on copies", covers every mode [17]. Each wrapper is copied once per occurrence in each list and never deduped by identity, so one object in two lists becomes two copies [17]. Deduping by identity would give two lists the same copy and bring the aliasing back [17]. I think this is the right fix for a library. The fusion functions receive state they do not own. One copy at the boundary covers all four modes with a single change.
The strongest check in the write-up is an ownership test. It sets the returned wrapper's score to 999.0, then asserts the cached score is still 0.4 [16]. The reviewer estimates the full audit at about 30 minutes per suspect function: grep for writes into inputs, assert on the caller's state, then test ownership [19].
What to watch
- A fix PR for the reciprocal_rerank half of #23351, and whether it copies wrappers at the dispatch sites or patches that one function.
- The first llama_index release that carries both #23333 and #23352.
- Whether other llama_index components that post-process NodeWithScore lists turn out to write into their inputs too, since the reviewer says the bug class generalizes far past this retriever.