Skip to content

Build1 publisher3 min readPublished

BM25 lost worst on the queries full of file paths. In 54% of them, the path is not in the answer

A CQADupStack benchmark reports dense retrieval beating BM25 by more on identifier-bearing queries than on the rest, because the identifier is often absent from the documents that answer them.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The benchmark uses BEIR / CQADupStack, the unix and android subforums: 70,380 documents and 1,771 judged queries.
  • The task is duplicate question detection: given a Stack Exchange question, find the earlier question that asks the same thing.
  • Before any retrieval runs, a classifier labels each query lexical or semantic; it only looks for lexical signals (something shaped like an identifier such as a file path, version number or error code, or text in backticks), and semantic is everything left over. The author notes there is no equally reliable way to detect that a query needs meaning-based matching.
  • 214 queries are labeled lexical and 1,557 semantic.
  • The BM25 side uses Porter stemming and identifier-aware tokenization.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A retrieval benchmark run on BEIR/CQADupStack, published by its author on dev.to, sorted every query into ones containing an identifier (a file path, a version number, an error code) and ones that did not, and found that the identifier queries were where BM25 trailed dense retrieval by the widest margin [1][3][19]. The stated reason is a corpus property you can measure before you buy anything: for 54% of those queries, the identifier appears in none of the documents judged to answer them [10].

The setup is 70,380 documents and 1,771 judged queries drawn from the unix and android subforums, with duplicate question detection as the task: given a Stack Exchange question, find the earlier question asking the same thing [1][2]. A classifier labels each query before any retrieval runs, looking only for identifier-shaped tokens or text in backticks and treating everything left over as semantic [3]. That yields 214 lexical and 1,557 semantic queries [4], so the lexical slice is about 12% of the workload [1], and the 54% figure covers roughly 116 queries [2].

The numbers, Recall@3 macro-averaged: BM25 with identifier-aware tokenization and Porter stemming scores 0.2255 overall, 0.3070 on lexical, 0.2143 on semantic [5][7]. The dense side, all-MiniLM-L6-v2, scores 0.3834, 0.4842 and 0.3695 [8]. So the gap is 0.1772 on the lexical slice against 0.1552 on the semantic one [3]. BM25 does do better on identifier queries than on the rest, by 0.0927 [4], but that is not a keyword advantage: the dense model gains even more on the same slice, 0.1147 [4]. Identifier queries are simply easier for both.

The mechanism the author gives is specific. BM25 can only match `xorg.conf` against documents that literally contain `xorg.conf`, and in a pool of duplicate questions most of those are somebody else's unrelated problem, while the true duplicate writes `/etc/X11/xorg.conf`, a different version number, or nothing at all [10][15]. A rare token also carries very high IDF, so on the occasions it does match it dominates the score and promotes a document whose only connection to the query is that one string [14]. The bi-encoder ignores the identifier and matches the problem description around it, which the duplicate does share [15].

On robustness: the BM25 arm was stemmed deliberately, since the Anserini Lucene analyzer behind the published BEIR reference numbers stems by default, with unstemmed results kept in the repo as a control [6]. The dense nDCG@10 of 0.4071 is close to the published reference for that model, which is the author's basis for trusting the harness [9]. A paired bootstrap over 10,000 resamples keeps the lexical interval clear of zero, and a sign test on the overall comparison returns 429 wins to 74 losses at p = 8.1e-62 [11][12], which leaves 1,268 queries tied [5]. Splitting the two subforums into separate datasets preserves the result: unix +0.2202 at p = 3.4e-07, android +0.1940 at p = 2.2e-05 [13].

The author is explicit about scope: duplicate question finding is close to the worst case for the assumption that identifiers are shared, because query and target are two people describing one problem, and in log or code search the identifier is shared by construction, where the usual advice should hold [16]. What breaks is the inference that a query containing keywords is a query helped by keyword matching [19]. That inference is the load-bearing part of routing traffic between a sparse and a dense arm [17].

Worth measuring before your next retrieval build: what share of your queries' identifiers actually appear in their own correct answers [18]. On this corpus that single ratio predicted the whole result [10][19].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories