Build1 publisher3 min readPublished
Plain-language questions push SEPA rulebook answers out of a top-5 vector search
One developer's test on 484 SEPA rulebook passages found that plain-English questions push several answers out of a top-5 vector search. Because the test measures each answer's rank directly, the failure shows up in retrieval, before the language model writes anything.
The Engineer · Build desk

What happened
- The live system passes only the five passages closest by cosine similarity, while deduplication and a Postgres full-text hybrid sit in the code switched off.
- Each test question was asked three ways: the expert original, a plain version barred from terms like PSP and SCT Inst, and a terse three-to-five-word search.
- The score was the labelled page's rank among all 484 passages, where rank 6 or worse means the model would never have seen the answer.
- The author's predictions that deduplication would help plain questions most and that a multi-word match rule would repair the hybrid search both failed.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A fixed top-5 cutoff shows the model about 1% of this corpus, so swapping in a stronger language model cannot recover an answer the retriever ranks sixth.
- decision A golden set written in the documents' own vocabulary overstates recall, so teams have to choose whether to add plain variants with a banned-term check before users find the gap.
- exposure Questions sitting at rank 4 or 5 under expert wording are the first to break for real users, and only a rank measurement shows how close to the cutoff they are.
- constraint The two switches already in this codebase did not close the gap as predicted, so turning them on is not a ready fix for plain-language queries.
Row 21 of the test set started as a bug report [5][15]. In an earlier piece, the author found a question whose answer sat at rank 7. The model only ever saw ranks 1 to 5, so it refused [2]. A reader, Mikhail, had found that question by hand. He then found two rewordings, both using the section's own words, that put the same passage at rank 1 [3]. His refused question is now the row's plain phrasing, and his working rewording is the original [15].
The plain column is built so the person writing the test cannot slide back into the rulebook's vocabulary [9]. It is the question as a customer or a new joiner would type it: "bank" for PSP, "instant transfer" for SCT Inst [9]. Each question has a list of forbidden terms, and the runner refuses to start if a plain phrasing contains one [9]. The phrasings were written before the run, and the golden set was not edited after the ranks came back [16]. The terse column, three to five words, may reuse section terms, because the variable there is length [9].
The corpus is the European Payments Council's two SEPA credit transfer rulebooks. They are cut into 484 passages of about 300 words each and embedded with OpenAI's text-embedding-3-small [4]. The live retriever hands the model the five passages closest by cosine similarity, and nothing else [6]. Five out of 484 is about 1% of the corpus [1]. If the passage holding the answer is not among them, the model cannot answer, however good it is [1]. So the score is the rank of the first labelled page among all 484, recorded once per phrasing against the stored production vectors [8][14]. Four questions already sat at rank 4 or 5 with their expert wording [10]. A judge grading final answers could pass all four [8]. Their ranks put them one or two slots from the cutoff [10].
The author tested four hypotheses. The fourth could not change anything, by construction, because every phrasing was English [11]. Of the rest, the author wrote: "The first held. The second and third did not." [12] The first predicted that plain wording would push several answers out of the top five, and terse wording more [11]. The second predicted that dropping duplicates would help plain and terse phrasings most, by freeing slots that duplicate passages take up [11]. There is a lot to drop. Of the 484 passages, 138 (about 29%) are exact duplicates, because a 36-page annex is bound into both rulebooks [7][2]. The third went after the hybrid path, which had cost 3 of 20 English questions in an earlier test [13]. The author blamed passages that matched on a single shared acronym [13]. The fix counted a passage only when it matched 2 or 3 distinct query words [14]. The excerpt ends at the results heading, so it does not include the per-question ranks or how far each answer fell.
The first result carries over to another system only if three conditions hold. Its users would have to ask questions in words the documents do not use. Its retriever would need a fixed top-k with no reranker behind it. Its embedder would need to handle that vocabulary gap about as text-embedding-3-small does. Payment rulebooks are a hard case on the first condition, since the plain phrasings had to replace terms like PSP and SCT Inst [9]. In my view, any golden set written by people who know the documents needs a plain column with a forbidden-term guard. Building one is cheap. The run used a read-only connection, one embedding per phrasing, and the production full-text query built on plainto_tsquery and ts_rank [14].
What to watch
- The per-question ranks for the plain and terse phrasings, which would show how many of the 21 answers fell past rank 5 and how far.
- A follow-up test of any fix for plain-language questions, given that deduplication and the multi-word hybrid rule both missed their predictions.
- A rerun on the same 484 passages with a different embedding model, to separate text-embedding-3-small's behaviour from the rulebooks' vocabulary.