Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

pg_search 0.26 trades single-term speed for a fourfold gain on ten-term BM25 search

ParadeDB's pg_search 0.26 cut a ten-term BM25 search from 129ms to 29ms in a benchmark published on dev.to. The same storage change made single-term queries several times slower, so teams whose searches are mostly one or two words would pay for the upgrade.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying pg_search 0.26 trades single-term speed for a fourfold gain on ten-term BM25 search
Generated illustration

What happened

  • On the ten-term OR query, EXPLAIN shows 0.25.11 running three scored sub-queries and merging them, where 0.26.0 runs a single one.
  • A single-term search for 'kernel' read 496 buffers on 0.26.0 against 132 on 0.25.11, and 382 of those reads were fieldnorms.
  • The author's runs at two and four terms place the point where 0.26.0 starts winning somewhere between two and four OR terms.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Whether to upgrade depends on how many OR terms a team's real queries carry, so the query logs need profiling before the version bump.
  • cost Teams serving short keyword lookups would pay roughly three to five times the latency on those queries in exchange for a multi-term gain they never use.
  • contradiction ParadeDB's claim that search got faster holds for long disjunctions in this third-party test and fails for single-term queries.
  • constraint Because the numbers come from a 260-word synthetic vocabulary, they say little about corpora with different term frequencies until someone reruns them.

pg_search adds a BM25 index type (USING bm25) to Postgres, so relevance-ranked search runs inside the database [2]. BM25 needs each document's field length to score it, and pg_search calls those numbers fieldnorms [3]. Up to 0.25.11 they sat in one shared array. In 0.26.0 they live in each term's postings list [3]. According to the post's author, every term's postings list now carries its own copy of the field-length data, and for a single term that is pure overhead [15].

The query plans show what the ten-term query gains. On 0.25.11, EXPLAIN lists Queries: 3 under the custom scan node. That means three scored sub-queries, merged afterwards. On 0.26.0 the same SQL runs as Queries: 1 [8]. The author credits the 4.3x drop to that collapse, with fieldnorms sitting next to the postings, and rules out cache warm-up [10]. The new plan also breaks buffer hits down by structure: 645 in total, 407 on postings, 190 on field norms [9]. That breakdown is new in 0.26.0 [9]. It is good engineering. Every page read is assigned to a structure, and that is how the single-term regression was diagnosed [14].

ParadeDB's own framing of the release is that search got faster [19], and the single-term case measured the other way. A query for 'kernel', a term in 19% of rows, ran roughly five times slower on 0.26.0 [11]. The rarer 'udp' went from 1.5-1.6ms to 4.8-5.7ms [12], a slowdown of 3.0x to 3.8x [22]. The author first suspected a broken harness, which is the right first suspicion; five reruns on each version gave the same split [13]. Buffer hits for 'kernel' rose from 132 to 496, and 382 of them were field norms [14]. That is 77% of the reads [21].

One figure fits the copy-per-term explanation less neatly. The single-term query read 382 fieldnorm buffers, about twice the 190 the ten-term query read [23]. The post does not explain why one term touches more field-length pages than ten. I'd want that answered before assuming the crossover holds on another table.

The author puts the crossover between two and four OR terms, with the gap widening after that [16]. "A query with one or two search terms gets no benefit from this release and pays a real cost; a query with four or more gets a win that keeps growing," the author wrote [17]. A four-term AND query moved from 8.7-10.7ms to 9.5-11.4ms, a regression inside the noise of three-run samples [18].

The test ran on one synthetic table: 3 million rows, with a title column drawn from about 260 words on a Zipfian distribution, so 'index' appears in 84% of rows [4]. The bm25 index used default options [5]. The author started from the view that "no rewrite is free" [20]. For the ten-term result to carry over, a team's traffic has to look like the benchmark query. That means an OR across many terms, ranked, with a LIMIT. The author says that is the shape ParadeDB's own benchmarking targets [7]. A workload dominated by one- and two-word searches sits on the slow side of every comparison in the post [11] [16].

What to watch

  • Whether ParadeDB publishes its own single-term numbers for 0.26 or offers a way to keep fieldnorms in a shared array.
  • A rerun on a real corpus with a larger vocabulary, to see if the two-to-four-term crossover holds.
  • An explanation of why a one-term query reads about twice as many fieldnorm buffers as a ten-term query.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    ParadeDB shipped pg_search 0.26.0 on 3 October; the changelog describes it as a rewrite of how the extension stores BM25 field-length data for scoring.

    ReportedSupportedSource: dev.to postView cited source
  2. [2]

    pg_search is a Postgres extension that adds a BM25 full-text index (USING bm25) so relevance-ranked search can run without leaving Postgres.

    ReportedSupportedSource: dev.to postView cited source
  3. [3]

    pg_search 0.26 moves fieldnorms, the per-document field-length numbers BM25 needs for scoring, from one shared array into the postings list of each term.

    ReportedSupportedSource: dev.to postView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 8, 2026

    ParadeDB's pg_search 0.26 cuts a ten-term BM25 search from 129ms to 29ms

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Full-text search in PostgresFollow
  • Database performance benchmarkingFollow
Loading related stories