Build1 distinct publisher3 min readUpdated
A dev.to argument on why benchmark charts mislead holds up on the mechanics: fame puts test sets into training data, and vendors pick which bars to show. The fix is an eval nobody outside your team has seen.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Every model launch arrives with a chart of bars or a spider diagram showing the new system edging past its rivals on benchmarks whose acronyms most people cannot expand [1]. A piece published on dev.to makes the case that this happens on a predictable schedule: within a week, users report that the new state-of-the-art model is, for their actual work, about the same as the last one or occasionally worse [2].
The mechanism worth understanding is contamination. Popular benchmarks are published and discussed on the open web, which is precisely where training data comes from [3]. When the questions and answers sit in the study material, a high score measures memorisation as much as ability, and nobody has to cheat deliberately for this to happen [4]. A model can score brilliantly on a test it has effectively already seen and then flounder on a genuinely novel version of the same task [5]. The dev.to framing is that fame is the failure mode: a benchmark stops measuring capability the moment it is well known enough to end up in the corpus [6].
Sitting on top of that is an incentive problem. A high score is not only an engineering result but a marketing asset worth a great deal in attention, funding and credibility [7]. Vendors choose which benchmarks to headline, which comparisons to draw, and which unflattering results to bury in an appendix or omit, which makes the launch slide a curated argument rather than a neutral readout [8]. The author is careful that this is rarely fraud; it is the ordinary gravity of a metric that has become a sales tool, with engineering effort flowing toward the number whether or not the number corresponds to anything you care about [9]. That is Goodhart's law in its plainest form: when a measure becomes a target it stops being a good measure [10]. A model tuned for benchmark-shaped questions is not necessarily better at your unglamorous, benchmark-unshaped problem [11].
For an operator the important point is directional. Contamination pushes observed scores above true ability, and selective reporting pushes presented scores above measured ones. Both errors have the same sign, so the leaderboard is not noisy around the truth, it is biased above it [18].
Even a clean, ungamed benchmark would mislead, because these tests favour what is cheap to score automatically: multiple choice, single verifiable answers, self-contained puzzles [12]. Your work is a long ambiguous document, a vague request, a task where good is a matter of taste and there is no answer key [13]. A model that aces graduate-level multiple choice can still write emails you would be embarrassed to send [14]. The properties you actually pay for, per the same piece, are largely unmeasured: precise instruction-following, consistent tone, not refusing sensible requests, coherence over a long session, and latency low enough not to break your flow [15].
Preference arenas, where humans vote between two anonymous answers, are more useful than static exams but measure preference rather than correctness [16]. Voters tend to prefer longer, more confident, more flattering replies, so a model can climb by being a better sycophant [17].
The conclusion is mine rather than the source's: the only eval with integrity is one that has never been published. Fifty to two hundred real tasks from your own backlog, graded by the person who owns the output, kept off the public internet. It is unglamorous and it does not survive being shared.
What to watch: whether vendors start publishing held-out or rotating test sets with contamination audits, and whether your own scores drift when you upgrade. If an internal set stops separating models, it has probably leaked or gone stale, and it needs replacing rather than defending.
Ranked by verification strength, evidence, and original report placement.
Every model launch comes with a chart, usually bars or a spider diagram, showing the new model edging past its rivals on a row of benchmarks with acronyms most people cannot expand.
Many popular benchmarks are published, discussed, and sitting on the open web, which is exactly where models get their training data.
When the questions and answers to the exam are in the study material, a high score measures memorisation as much as ability; nobody needs to cheat deliberately because the leak is structural.
A model can score brilliantly on a benchmark it has effectively already seen and then flounder on a genuinely novel version of the same task.
A benchmark stops measuring intelligence the moment it becomes famous enough to end up in the training data; fame is the thing that breaks it.
A high benchmark score is not just an engineering result but a marketing asset worth an enormous amount in attention, funding and credibility.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 15, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Mechanism argument, no measurement
One self-published essay carries the entire cluster. Its core mechanisms — public benchmarks overlapping training corpora, Goodhart targeting of a scoreboard, auto-scorable task shapes diverging from real work, preference rather than correctness in arenas — are internally coherent and stated clearly, which is why several claims are marked supported. But nothing is instrumented: no benchmark, model, vendor, contamination study, or arena dataset is named, and the empirical generalisations about post-launch user experience and human preference bias have no data behind them at all.
No adoption signal supplied
The cluster contains no release, deployment, pricing, licence, benchmark-run, or usage disclosure. The essay does not report anyone adopting private held-out evals, nor any vendor changing benchmark reporting practice, so there is nothing to measure and no adoption observation can be recorded without inventing one.
Sweeping conclusions on thin substantiation
The direction of the argument is deflationary — it discounts vendor hype rather than amplifying it — and the contamination and incentive mechanics are plausible, so the gap is modest. It is positive rather than zero because the piece generalises to all popular benchmarks and to a directional upward bias in published figures without naming a single contaminated benchmark, quantifying any score inflation, or acknowledging decontamination and held-out-set practices that would bound the effect. The prescription that your eval should be one nobody has published is asserted as the fix without any evidence that private evals produce better decisions.
Vendor scorekeeping incentives central; author incentives light
The cluster's own subject is an incentive structure: the source states that a high score is a marketing, funding, credibility, recruiting and press asset, that a one-point gain can convert into headlines and a valuation bump, and that vendors therefore select which benchmarks to headline and which results to omit. Those incentives are asserted by one commentator rather than documented, so the score reflects a strong and specific incentive thesis with weak substantiation. On the publishing side the exposure is mild: a dev.to author with a critical stance and no disclosed vendor relationship, though a contrarian evaluation-critique post also carries its own attention incentive.
Single-publisher argument, uncorroborated
Confidence is limited by structure rather than by internal quality: one publisher, one article, no independent corroboration, and no adoption dimension at all. The conceptual claims (contamination mechanism, Goodhart targeting, preference versus correctness) are the kind that hold up on reasoning and are unlikely to be reversed in direction, which keeps the score above floor; the empirical and magnitude claims could easily be overturned by a single measured study.
Follow any of these and your For You feed starts watching them — no settings page required.
build
A docs bot that refuses to answer is working: the case for an evidence gate over a bigger window1 distinct publisher
build
Your token ratio, not the leaderboard, decides which model is cheap1 distinct publisher
build
Force the tool call, then hand Lightsail a long-lived key1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher