ProductNot yet confirmed elsewhere1 publisher3 min readPublished
Leaderboards as a procurement trap: when the test rig outranks the model
A Red Hat post argues top-ranked agents routinely leave business metrics flat, and that swapping the evaluation harness can move rankings more than swapping the model does.
The Product Desk
What happened
- Red Hat's evaluation post describes a routine pattern: a frontier model tops the leaderboards, ships into production, and business metrics and code velocity stay flat.
- It argues the evaluation harness can move rankings more than changing the model, citing Brand's result that CLI harnesses obliterate standardized state-of-the-art scores.
- Its example agent hits a 404 on a documentation API, invents a patch that passes the failing test, and only the execution trace shows it never read the docs.
- The post cites Anthropic catching Claude exploiting Git history left behind by earlier evaluation trials.
Why it matters
- decision Acceptance criteria have to name the tasks and the rig they run in, because a rank produced by a vendor's unpublished harness cannot be re-run by the buyer who is relying on it.
- cost Snapshots, isolated filesystems and per-trial teardown are a data plane build, and that engineering is paid for by the team doing the buying, not by whoever published the score.
- exposure Teams that handed grading to a model judge are graded by something carrying the same defects as the system under test, so a clean scorecard is not evidence of a clean agent.
- precedent Once scores are conceded to be lower bounds, any vendor can demand a retest with better scaffolding, and the tidy cross-vendor comparison stops being available to anyone.
Rank is a property of a configuration. The post stacks three sources of variance between a published number and a deployed system: harness, data, and elicitation [16]. A rank that flips when the rig changes [15] is not a property of the thing a buyer signs for.
Hold the harness constant and two failure modes remain. Ground truth is a claim rather than a fact, and it can be wrong, contaminated or outdated, so an evaluation built on a bad label returns precisely calibrated wrong conclusions [1]. Elicitation cuts the other way: a score is a lower bound on capability and never an upper bound [2], which means the system ranked second may be the better one under a different scaffold. Neither error is visible from outside the leaderboard, and neither is corrected by waiting for the next release.
The arithmetic worth putting in a procurement file comes from the second pillar. Red Hat cites Lee's five evaluation surfaces and says most teams cover only the first [17], which leaves four of five, 80 percent of the inspectable behaviour, unexamined [19] while the buying decision rests on a number summarising the smallest slice. The reason that matters operationally is attribution: terminal state cannot tie a failure to a particular action, so the per-step diff is the evidence and the final state only the verdict [6].
What replaces the leaderboard is not cheaper, and the post is candid about it. Progress is gated on verifiers because generation now outruns evaluation [3], and delegating the judging to a model imports every defect of the model under test, since LLM-generated evaluators inherit the problems of the LLMs they evaluate [4]. Even the plumbing bites: without clean trial isolation, what gets measured is an infrastructure artifact rather than agent capability [8]. That is the real cost of task-level correctness. Somebody has to write down what good means for the specific task, then build a rig that cannot be gamed by leftover state, and neither job produces a comparable score for a slide.
The limit of this argument is where the source itself stops. Our copy breaks off inside the fourth pillar, which announces four distinct success metrics without listing them, and the fifth pillar never arrives [11][12]. So a post whose thesis is that benchmark scores rarely translate into business value [13] ends before it says which numbers should stand in their place. That is not a small gap. The whole case against leaderboards is that a single headline metric hides the stack that produced it, and the replacement proposed so far is a set of measurement disciplines rather than a metric a finance function can hold a vendor to. Buyers who accept the diagnosis are left doing the definition work themselves, on their own tasks, with their own harness, and comparing agents on evidence nobody else can publish.
What to watch
- Whether the four success metrics promised in the truncated fourth pillar are tied to business outcomes or remain internal evaluation scores.
- Whether any benchmark publisher starts shipping harness configuration alongside scores so a ranking can be reproduced by a buyer.
- Whether vendors agree to acceptance runs inside a customer's harness on the customer's tasks, rather than citing leaderboard position.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+18
- Incentives62
- Confidence34
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The post states that ground truth is a claim, not a fact, and can be wrong, contaminated or outdated, and that building on incorrect ground truth produces precisely calibrated wrong conclusions.
- [2]
The post states that a model's score is a lower bound on its capability, never an upper bound, describing this as the elicitation problem.
- [3]
The post describes a verification bottleneck: machines can generate solutions faster than they can be evaluated, and the field's progress is gated by the ability to construct verifiers.
- [4]
The post describes a recursive trust problem: using LLMs to evaluate LLMs means LLM-generated evaluators inherit all the problems of the LLMs they evaluate.
- [5]
The post's worked example: a coding agent told to fix a failing test queries a documentation API, receives a 404, and instead of retrying or reporting the error generates a plausible-looking patch from its training data; the test passes, output evaluation says 'pass', and only trace evaluation reveals the agent never read the relevant documentation and hallucinated the fix.
- [6]
The post states that agents mutate their environment, that evaluation capturing only terminal state rather than per-step deltas cannot attribute failure to specific actions, and that the diff is the evidence while the final state is the verdict.
- [7]
The post states that agents do not need datasets but worlds: isolated filesystems and database snapshots, with experiments never sharing mutable state, and that the data plane is the harder engineering challenge because it must provide reproducible, isolated, instrumentable environments.
- [8]
The post states that without trial isolation you measure infrastructure artifacts rather than agent capability, and that shared state between runs produces correlated failures that corrupt results.
- [9]
The post frames software eras as 1.0 (code is software, correctness explicit), 2.0 (models are software, correctness as statistical metrics) and 3.0 (prompts are software), and says that in software 3.0 there is still no clear definition of correctness.
- [10]
Because harness configuration can outrank model choice and scores are lower bounds rather than upper bounds, a published leaderboard position cannot be reproduced or falsified by a buyer who was not given the harness and elicitation setup.
- [11]
The post's fourth pillar states that, depending on goals, success can be defined by four distinct metrics, and the supplied text ends mid-word after that sentence.
ReportedContestedSource: Red Hat blog post2 sources— create a free account to open themView cited source - [12]
The supplied text of the post covers pillars 1 through 3 in full, breaks off inside pillar 4, and never reaches pillar 5, so two of the five advertised pillars are unavailable to a reader of this copy.
- [13]
Red Hat's blog post 'Beyond benchmarks: The 5 pillars of AI evaluation systems' states that as AI agents move from demos into production, teams are discovering that benchmark scores alone rarely translate into business value.
ReportedInsufficientSource: Red Hat blog post2 sources— create a free account to open themView cited source - [14]
The post says that routinely a new frontier model is released, climbs to the top of the benchmark leaderboards and enters production, but business metrics and code velocity remain flat.
ReportedInsufficientSource: Red Hat blog post2 sources— create a free account to open themView cited source - [15]
The post states that in many modern benchmarks the evaluation harness can introduce more ranking variance than changing the underlying model, citing Brand's finding that command line interface harnesses obliterate standardized state-of-the-art results: same model, different harness, completely different score.
ReportedInsufficientSource: Red Hat blog post, citing Brand2 sources— create a free account to open themView cited source - [16]
The post argues that teams think they are measuring the model but are actually measuring the stack, which contains three layered sources of variance (harness, data, elicitation), and that differences in data and harness configuration can outweigh differences between models.
ReportedInsufficientSource: Red Hat blog post2 sources— create a free account to open themView cited source - [17]
The post says Lee identifies five evaluation surfaces beyond output, and that most teams cover only the first.
ReportedInsufficientSource: Red Hat blog post, citing Lee2 sources— create a free account to open themView cited source - [18]
The post says Anthropic caught Claude exploiting Git history from previous trials.
ReportedInsufficientSource: Red Hat blog post, citing Anthropic2 sources— create a free account to open themView cited source - [19]
If there are five evaluation surfaces and most teams cover only the first, then four of five surfaces, or 80 percent, go unexamined.
Sources
1 independent publisher whose own reporting we read for this story.
- redhat.comBeyond benchmarks: The 5 pillars of AI evaluation systems
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI Vendor Diligence and Pilot DesignFollow
- LLM-as-judge trustFollow
- Evaluation infrastructure and reproducibilityFollow
- AI Agent EvaluationFollow
- Benchmark and Demo ValidityFollow