Skip to content

Build1 publisher3 min readPublished

The 21-cent model bake-off that inverted when the judge got audited

Five models, ten questions, a tidy leaderboard. Then the author checked who was grading, found a contestant holding the pen, and re-scored the saved answers for three cents.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying The 21-cent model bake-off that inverted when the judge got audited
Generated illustration

What happened

  • The author built a pytest harness that runs the same set of questions through five language models at once: a free local Llama, plus GPT, DeepSeek, and two Claude models, comparing cost per query, speed, and answer quality.
  • The run design was five models, the same ten questions, twice each, with every call measured.
  • It is the first project in the series that spends real money; the whole comparison came to about 21 cents.
  • Quality was scored by a second model grading each answer on correctness and relevance, combined into a 0-1 score with a pass line of 0.7.
  • deepseek-v4-pro: quality mean 0.970, $0.000138 per query, 2713 ms mean latency, 113 output tokens.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A pytest harness that pushed the same ten questions through five language models, twice each, measuring cost, latency and a graded quality score, produced a clean leaderboard for about 21 cents [1][2][3]. Its author, writing at dev.to, then audited the thing doing the grading, found it was one of the five contestants, and re-scored the saved answers with a paid judge [14][15].

The first scoreboard, ordered by quality: deepseek-v4-pro at 0.970, $0.000138 per query, 2713 ms [5]; claude-haiku-4-5 at 0.967, $0.000537, 1597 ms [6]; gpt-5.6-luna at 0.962, $0.000082, 1323 ms and 65 output tokens [7]; claude-sonnet-5 at 0.937, $0.002426, 4093 ms and 239 output tokens [8]; and local llama3.2 at 0.922, no marginal cost, 7859 ms [9]. The entire quality column sits between 0.92 and 0.97 [10]. Sonnet runs about 30 times GPT's per-query cost [1] and scores no higher, its answers merely longer [11]. The free local model is roughly six times slower than GPT on latency [3].

Every quality number came from a second model scoring correctness and relevance into a 0-1 value with a 0.7 pass line [4]. Before ranking anything, the author put a range around each score and around each pairwise difference on the same questions; every paid-versus-paid interval crossed zero, which makes the paid models a tie on this question set and the order between them noise [12]. Supporting that: GPT placed third in this run and first in an earlier run of the same setup [13].

The audit is the part worth copying. The judge was the free local llama3.2, which is also one of the five models being graded [14]. Re-grading the 200 saved answers with a paid judge required no new model calls, since the judge only reads text already on disk, and cost about three cents, roughly a seventh of the original run [15][2]. Under the paid judge, haiku, GPT and DeepSeek all came in at 1.000 [16][17][18], Sonnet at 0.997 [19], and llama3.2 went the other way, from 0.936 down to 0.919 [20]. Pass rates tell it more plainly: every paid model held at 40 of 40 while the local model fell from 39 of 40 to 35 of 40 [21], 97.5 per cent to 87.5 per cent, four additional failures [5]. Those denominators imply 40 graded answers per model across the five, the 200 that were re-scored [4].

According to the author, the cheap judge had been closing a gap that the better judge shows is real [24]. Reading the cases where the two judges disagreed, the disagreements were not random: the cheap judge was consistently too soft on the weak model's wrong answers [22]. And those wrong answers only existed because of a second, harder question set, written after the easy ten failed to separate anything and everyone passed [23].

Two things to watch. First, whether an ordering built on a non-contestant judge survives repetition, given that the paid tie held across runs while the ranking did not [12][13]. Second, the cost asymmetry: generating answers cost 21 cents and re-grading them cost three [3][2], so a grader audit is the cheapest experiment in the harness and there is no budget argument for skipping it.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories