Skip to content

Build1 publisher3 min readPublished

Crutches built from measured failures lift a local Qwen 3B from 33% to 52% on post-cutoff facts

Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Crutches built from measured failures lift a local Qwen 3B from 33% to 52% on post-cutoff facts
Generated illustration

What happened

  • The new 150-question set drops every question the bare 3B can answer unaided, and on what is left the bare model scores 0.0%.
  • The stack runs fully locally on a 49 GB FTS5 index of 7.2 million Wikipedia pages and 11.4 million redirects.
  • Swapping in an untuned 4-bit Qwen2.5-7B as reader, with the rest of the chain unchanged, gave the larger model a 16-point lead on the primary arm.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Every promotion was decided on the same 150 questions, so 52% is a best case for that set, and fresh questions of the same type should be expected to score lower.
  • cost In these runs, five-way voting multiplied inference cost and bought nothing, because a wrong page yields the same wrong answer on every sample.
  • decision Whether to tune a small reader or buy a bigger one depends on failure modes, since an entropy gate calibrated on the 3B does nothing for a 7B sitting near 0.002.
  • capability With the set public on Kaggle, other stacks can be scored for net retrieval value, though the zero floor has been established only for the 3B.

The contamination case rests on one measurement. With no retrieval, the bare 3B answered 30.5% of HotpotQA, and the author puts that score down to memorization [1]. The wider claim, that the 2018 dataset sits in the training data of every 2024 model, is the author's assertion [2]. On a benchmark like that, the author argues, retrieved evidence can only hurt [3].

The post-cutoff-150 set is built to strip out the memorized share. Its questions come from Wikipedia pages whose leads cite 2025 and 2026 facts, and each answer is an exact span of the page [4]. Any question the bare 3B could answer without context was rejected. On the remainder the bare model scores 0.0%, with a Wilson upper bound of 2.5% [5]. The filter was defined against the 3B, so that floor has been measured for the 3B alone [5].

Every change then faced one gate. A challenger replaced the champion only if it beat it on the same 150 questions through the same matcher, and rollbacks were published alongside wins [9]. Publishing the rollbacks is good practice. But promoting challenger after challenger against one fixed set means the final configuration was chosen on the test set, and the post does not describe a held-out split. At n = 150, the Wilson 95% interval around 52% runs from about 44% to 60% [1].

The 7B run changed only the reader, to a 4-bit Qwen2.5-7B-Instruct with no LoRA and no tuning. It won the primary arm by 16 points [16]. The crutches then added 28 points to the 3B and 9 to the 7B [17]. If both gains count from each reader's primary-arm score, the 3B finishes about 3 points ahead [2]. The author concedes the intervals overlap, claims a tie or better, and says a tuned 7B would do better [20]. The 3B's entropy runs 0.05 to 0.21 when it is right and above 1.0 when it guesses. The 7B's sits near 0.002, so the calibrated gate is useless on it, and it has no echo for the SLOTS fix to remove [18]. "The crutches don't transfer because they were never generic," the author wrote [19].

In my view the rollback list is the most useful part of the post. Five-way self-consistency returned nothing, as did best-of-N by confidence, because the errors were systematic: a wrong page gave the same wrong answer five times [12]. IRCoT scored 10% naive and 25% hybrid, at least 17 points behind the sequential chain [13]. SpanLift reached 1.3% [14]. The author fixed GRPO's sparse rewards, cutting frac_reward_zero_std from 0.8 to 0.2, and the resulting adapters still scored the same as the champion [15]. DPO was the only technique that broke a plateau, and it did it twice [8].

The infrastructure rules are good engineering. Evaluating a model in memory straight after training measured 0% on a healthy model, so the harness saves, reloads from disk and only then gates [10]. A zero is at least a loud bug. Temperature 0.1 on extraction calls was quieter, adding 4 to 6 points of run-to-run variance to every gate until all extraction moved to temperature 0 [11]. At the top of that range, the noise is two-thirds of what the crutches gave the 7B [3].

For the 52% to carry over, the questions have to look like these: a single answer, present verbatim in a Wikipedia lead and reachable by anchoring on a page title [4][13]. The reader also has to fail the way this 3B fails, because each crutch was built on a measured 3B failure mode [22]. The author draws the same line from the IRCoT result: for a 3B, anchoring by exact titles beats conversational retrieval [13].

What to watch

  • A tuned 7B run through the same chain, which the author names as the next experiment; a clear lead over 52% would limit the small-model result to zero-shot comparisons.
  • Naked scores for other models on the Kaggle post-cutoff-150 set, showing whether a filter defined against the 3B still yields a zero floor for them.
  • A fresh question set built the same way and scored with no further gating, to show how much of the 52% survives outside the 150 questions it was tuned on.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories