Build1 publisher3 min readPublished
Benchmarks are contaminated by design: your eval set should be one nobody has published
A dev.to argument on why benchmark charts mislead holds up on the mechanics: fame puts test sets into training data, and vendors pick which bars to show. The fix is an eval nobody outside your team has seen.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Every model launch comes with a chart, usually bars or a spider diagram, showing the new model edging past its rivals on a row of benchmarks with acronyms most people cannot expand.
- Within a week of a launch, users report that the new state-of-the-art model is, for their actual work, about the same as the last one or occasionally worse.
- Many popular benchmarks are published, discussed, and sitting on the open web, which is exactly where models get their training data.
- When the questions and answers to the exam are in the study material, a high score measures memorisation as much as ability; nobody needs to cheat deliberately because the leak is structural.
- A model can score brilliantly on a benchmark it has effectively already seen and then flounder on a genuinely novel version of the same task.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Every model launch arrives with a chart of bars or a spider diagram showing the new system edging past its rivals on benchmarks whose acronyms most people cannot expand [1]. A piece published on dev.to makes the case that this happens on a predictable schedule: within a week, users report that the new state-of-the-art model is, for their actual work, about the same as the last one or occasionally worse [2].
The mechanism worth understanding is contamination. Popular benchmarks are published and discussed on the open web, which is precisely where training data comes from [3]. When the questions and answers sit in the study material, a high score measures memorisation as much as ability, and nobody has to cheat deliberately for this to happen [4]. A model can score brilliantly on a test it has effectively already seen and then flounder on a genuinely novel version of the same task [5]. The dev.to framing is that fame is the failure mode: a benchmark stops measuring capability the moment it is well known enough to end up in the corpus [6].
Sitting on top of that is an incentive problem. A high score is not only an engineering result but a marketing asset worth a great deal in attention, funding and credibility [7]. Vendors choose which benchmarks to headline, which comparisons to draw, and which unflattering results to bury in an appendix or omit, which makes the launch slide a curated argument rather than a neutral readout [8]. The author is careful that this is rarely fraud; it is the ordinary gravity of a metric that has become a sales tool, with engineering effort flowing toward the number whether or not the number corresponds to anything you care about [9]. That is Goodhart's law in its plainest form: when a measure becomes a target it stops being a good measure [10]. A model tuned for benchmark-shaped questions is not necessarily better at your unglamorous, benchmark-unshaped problem [11].
For an operator the important point is directional. Contamination pushes observed scores above true ability, and selective reporting pushes presented scores above measured ones. Both errors have the same sign, so the leaderboard is not noisy around the truth, it is biased above it [18].
Even a clean, ungamed benchmark would mislead, because these tests favour what is cheap to score automatically: multiple choice, single verifiable answers, self-contained puzzles [12]. Your work is a long ambiguous document, a vague request, a task where good is a matter of taste and there is no answer key [13]. A model that aces graduate-level multiple choice can still write emails you would be embarrassed to send [14]. The properties you actually pay for, per the same piece, are largely unmeasured: precise instruction-following, consistent tone, not refusing sensible requests, coherence over a long session, and latency low enough not to break your flow [15].
Preference arenas, where humans vote between two anonymous answers, are more useful than static exams but measure preference rather than correctness [16]. Voters tend to prefer longer, more confident, more flattering replies, so a model can climb by being a better sycophant [17].
The conclusion is mine rather than the source's: the only eval with integrity is one that has never been published. Fifty to two hundred real tasks from your own backlog, graded by the person who owns the output, kept off the public internet. It is unglamorous and it does not survive being shared.
What to watch: whether vendors start publishing held-out or rotating test sets with contamination audits, and whether your own scores drift when you upgrade. If an internal set stops separating models, it has probably leaked or gone stale, and it needs replacing rather than defending.