Published Leadership3 min read
A vendor gamed a leaderboard in an afternoon. Price your benchmarks accordingly.
Speechmatics ran a hackathon to see how far benchmark gaming would get it and climbed 10 places on one retrain.
Context for builders, not their beat.See today for builders

What happened
- Katy Wigdahl is CEO of Speechmatics, a voice AI company, and wrote a Forbes Tech Council piece titled 'Are You Being Benchmaxxed?' published on forbes.com.
- Speechmatics ran an internal hackathon earlier this year whose brief was to game an AI benchmark the same way vendors do; the team climbed 10 places with only a few hours of work and one retrain, placing alongside models from other participant companies.
- The problem is structural: the moment a test set becomes the standard measure of quality, it also becomes the optimisation target, and optimising for a benchmark and improving a product have become two different activities.
- Teams fine-tune on data that resembles the test conditions, select inference settings tuned for evaluation, and often train on the test data itself, a practice purists frown upon.
- A 2025 study found that giving developers even limited additional access to test data could boost leaderboard scores by up to 112%.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
Speechmatics ran an internal hackathon with the explicit brief of gaming an AI benchmark the way vendors do, and climbed 10 places with a few hours of work and a single retrain, according to chief executive Katy Wigdahl writing for the Forbes Tech Council [1][2]. That matters because leaderboard position is still an input to procurement, and a ranking that moves that far, that fast, without anyone trying to improve the product is not measuring what buyers think it measures [3][9].
The mechanics are mundane. Wigdahl describes teams fine-tuning on data that resembles the test conditions, picking inference settings tuned for evaluation, and in some cases training on the test data itself [4]. She cites a 2025 study finding that even limited additional access to test data could lift leaderboard scores by up to 112%, and a February 2026 paper finding that nearly half of widely used benchmarks have saturated, with top models scoring so similarly that the tests no longer separate them [5][6]. Her framing of the incentive is the right one: vendors are not being dishonest, they are optimising rationally for the metric that drives purchase decisions [9].
The part of the hackathon worth stealing is the failure mode, not the climb. When the team fine-tuned further, error rates more than doubled at a later checkpoint, so the submission would have looked fine while a live deployment would not [7]. Rank and reliability moved in opposite directions inside the same experiment [11]. Separately, the exercise surfaced a feature that existed in the product, was configured, and was simply not connected to anything, producing no effect at all until someone fixed it [8]. Stress-testing the pipeline end to end found that; the spec sheet did not, and no benchmark would have [8].
Note the interest. This is a voice AI vendor's chief executive arguing that the public scores buyers use to compare voice AI vendors are unreliable [1]. The argument survives the conflict, because the supporting detail is an admission against interest: her own team demonstrated the gap, and she reports that an entire category of companies now exists to evaluate AI systems independently, a market she says barely existed three years ago [2][10]. Someone is monetising the trust deficit either way.
The structural point for operators is that the question has changed. Wigdahl argues models are now competing to be the best for a specific use case rather than best in the abstract, with speed, accuracy, cost, reasoning, multilingual performance and vertical domain expertise all optimised differently by different vendors [12]. The headline number is usually the score a vendor is most confident winning, and it tends not to disclose which user populations were excluded, which test conditions were left out, or which use cases never entered the evaluation set [13]. A contact centre with regional accents, background noise and overlapping speakers looks nothing like a vendor demo [14]. Her conclusion is that the advantage comes from internal discipline to test against your own conditions rather than tracking leaderboards [15].
Watch whether your next AI contract carries acceptance criteria measured on your data and your traffic, with a re-test clause at each vendor model update, or whether it cites a public score. Watch the saturation figure: if half of widely used benchmarks already fail to separate top models, the vendor-supplied comparison is decoration [6].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Katy Wigdahl is CEO of Speechmatics, a voice AI company, and wrote a Forbes Tech Council piece titled 'Are You Being Benchmaxxed?' published on forbes.com.
- [2]
Speechmatics ran an internal hackathon earlier this year whose brief was to game an AI benchmark the same way vendors do; the team climbed 10 places with only a few hours of work and one retrain, placing alongside models from other participant companies.
- [3]
The problem is structural: the moment a test set becomes the standard measure of quality, it also becomes the optimisation target, and optimising for a benchmark and improving a product have become two different activities.
- [4]
Teams fine-tune on data that resembles the test conditions, select inference settings tuned for evaluation, and often train on the test data itself, a practice purists frown upon.
- [5]
A 2025 study found that giving developers even limited additional access to test data could boost leaderboard scores by up to 112%.
- [6]
A February 2026 paper found that nearly half of widely used benchmarks have hit saturation, meaning top models score so similarly that the tests can no longer meaningfully separate them.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- forbes.comKaty Wigdahl, Forbes Councils MemberAug 13Are You Being Benchmaxxed?
Additional citations
- Forbes Tech Council byline for Katy Wigdahl, CEO of Speechmatics
- Katy Wigdahl, Speechmatics, in Forbes
- 2025 study cited by Katy Wigdahl in Forbes
- February 2026 paper cited by Katy Wigdahl in Forbes


