Skip to content

Build1 publisher2 min readPublished

Epoch AI's benchmark audit fails nine of the first 15 versions it inspected

Broken answer keys and graders that punish correct tool calls drove the verdicts. Epoch AI says it stops each review once it has enough evidence, so the published defect counts are floors.

The Engineer · Build desk

Illustration accompanying Epoch AI's benchmark audit fails nine of the first 15 versions it inspected

What happened

  • Epoch AI launched Benchmark Reviews on September 17th and labelled nine of the first 15 benchmark versions it audited as Flawed, four as Verified, and two as Not Enough Info.
  • Its published methodology gives a Flawed verdict when at least 20 percent of an inspected sample contains errors, or when a single issue corrupts grading at scale.
  • The Humanity's Last Exam review found substantial accuracy-altering errors in 22 of 48 sampled questions, and judged 12 of those questions impossible to answer correctly as written.
  • Berkeley Function Calling Leaderboard v4, which scores whether a model picks the right tool and passes the right arguments, showed potential accuracy defects in 24 of 50 sampled tasks.
  • DeepSWE v1.1 crossed the threshold narrowly, with confirmed false negatives in at least 23 of 113 tasks after the grader discarded changes an agent had made.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A team that put a Humanity's Last Exam or Berkeley Function Calling Leaderboard v4 number into a deployment or procurement case now has to defend a score with a public task-level defect list behind it.
  • constraint Because Epoch AI halts each review at the first sufficient evidence, a Flawed verdict gives a buyer the floor of the problem and leaves the size of the correction to whoever needs the number.
  • decision Version strings become part of the citation, since the audit treats tests that change without a version update as a defect in its own right.
  • capability A maintainer who disputes a Flawed label can rerun the same inspection at the same task counts, because the sampling and expansion thresholds are published.

The defects sit in answer keys and graders. Epoch AI's review of Berkeley Function Calling Leaderboard v4, which tests whether a model selects the right tool and supplies the right arguments [12], found a task that expected the model to add a stock already present on the watchlist [13]. A model that checks the list first and declines the duplicate is scored as a failure. Another task's web-search answer had become stale since the key was written [13].

In Humanity's Last Exam, one question asked for a four-point discrete Fourier transform and supplied eight entries [10]. Another answer key pointed at the wrong multiple-choice option [10]. Epoch AI reported problems that could reject valid answers and problems that could accept incorrect ones [11].

The sampling rules explain the sample sizes. Epoch AI takes 50 tasks from a benchmark with more than 50 available, expands to 100 when the observed error rate falls between 15 and 25 percent, and inspects every task when there are 50 or fewer [6]. DeepSWE v1.1's confirmed false negative rate is 20.3 percent [14], inside that band, and the 113 tasks inspected are past the expansion level [4]. Humanity's Last Exam at 46 percent and Berkeley Function Calling Leaderboard v4 at 48 percent are both above the band [8][2], so a 50-task sample settles them [5]. The Humanity's Last Exam sample was 48, two short of the stated 50, which at that error rate does not change the verdict [8].

Epoch AI stops a review once it has gathered enough evidence for a Flawed verdict, and says the writeups should be read as a minimum case [7]. A Flawed label bounds the defect count from below.

Nine Flawed out of 15 is 60 percent [1]. The registry Epoch AI maintains holds 85 benchmarks [15], so the first batch covers at most about a fifth of it [3], and the launch material does not say how the 15 versions were picked. For that 60 percent to transfer to whatever benchmark you are about to cite in a deployment review, the selection would have to have been blind to which tests already looked suspect. The per-benchmark findings hold on their own.

The reviews were completed between August 10th and September 12th [16], so the September 17th launch published an accumulated program. Epoch AI, co-founded and led by Jaime Sevilla, argues that AI progress needs better measurement than leaderboards and launch-day claims supply [18]. Epoch AI says its founding group was frustrated that a consequential industry was still being measured through "hype and vibes" [17].

What to watch

  • Whether Epoch AI publishes a selection rule for the next batch, and whether the Flawed share holds outside the first 15 versions.
  • Whether maintainers reissue corrected versions after a Flawed verdict, and whether Epoch AI re-reviews them under a new version number.
  • Whether model developers keep citing Humanity's Last Exam and Berkeley Function Calling Leaderboard v4 scores now that the sampled defects are public.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories