Skip to content

Build1 publisher2 min readPublished

Re-running five Terminal-Bench-Science tasks at $12 each leaves Fable 5.1 passing one

The published doubling was measured with a per-task allowance most teams will never grant, under a harness Anthropic has not described. The number that transfers is the one you get at your own spend cap.

The Engineer · Build desk

Illustration accompanying Re-running five Terminal-Bench-Science tasks at $12 each leaves Fable 5.1 passing one

What happened

  • Anthropic built the Claude Fable 5.1 launch around Terminal-Bench-Science, where its own scoring puts the new model at 52.6 percent against Fable 5's 24.7 percent, more than double.
  • The benchmark lets each model work up to eight hours on a task, and Anthropic has not said which harness or what budget produced its published score.
  • When the benchmark's own leaderboard ran Fable 5, it used Claude Code at maximum effort and spent $14,180 across 210 attempts, about $67 per attempt.
  • The New Stack re-ran five of the 70 public tasks, one per science field, giving each run a plain terminal, a $12 ceiling and 60 turns; the ten runs took about 12 hours.
  • Only the mathematics task produced a pass, with Fable 5.1 finding the hidden structure and stopping on its own, while Fable 5 failed the same task on two separate attempts.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Sizing an upgrade on the published delta means buying a harness you cannot inspect, so the only defensible acceptance test is your own turn and dollar cap on tasks you already run.
  • cost A true replication of leaderboard conditions is a five-figure evaluation line item, which prices most teams out of it and pushes them toward small samples that cannot settle the question.
  • contradiction One reading says the new model does twice the work of the old; at a consumer cap both fail four of the same five tasks, and those two readings support different procurement decisions.
  • constraint With all-or-nothing grading, a spend cap that stops a run before the results file exists is recorded as a capability failure, which makes cheap evaluations systematically pessimistic.

At $12 and 60 turns, Fable 5 ended four of five tasks by running out of money; Fable 5.1 never reached the cost ceiling, stopping instead at the turn limit or on its own [4]. That makes the comparison at this budget partly a measurement of cost per turn of progress. Symbolic regression separates the two effects: Fable 5.1 found the hidden structure in 27 turns and 27,088 output tokens, roughly 31 percent fewer than Fable 5 burned failing the same task on its first attempt [7][8][6].

The leaderboard's per-attempt spend on Fable 5 is about 5.6 times the cap used in this re-run [1]. That spend was the leaderboard's own choice: Anthropic has published neither the harness nor the budget behind 52.6 percent [3][4]. Time was not the binding constraint either: eight hours per task is allowed, and the longest single run took 139.3 minutes, about 29 percent of the allowance [3][5].

Caps truncate different failures. Both models spent the nanoindentation run reading raw curves and writing code to segment them, and neither produced a results file the grader accepted [13]. Fable 5.1 got further on the foraging task, building a model, testing it against its own scoring loop and declaring itself done at 43 turns for $5.65, before the official grader rejected it [12]. Self-termination looked efficient on the invoice, but the grader still scored it zero.

For 52.6 percent to transfer, your harness has to let a task run toward the full eight hours, neither turns nor dollars can bind before the model writes the artifact the grader wants, and you have to accept all-or-nothing grading on hidden data, which on the Lorenz-96 task means five criteria pass together or not at all [3][14]. Miss the middle condition and the score you record is a reading of your budget.

The re-run's own limits are worth stating: five of 70 tasks, chosen because they fit a Python environment, one run each apart from the repeated symbolic regression attempt [5][8]. One pass against none is too small a sample to count as a rate [3]. The New Stack also reports that on everyday work the two models sat far closer together than the benchmark implies [15]. Fable 5.1's five runs cost $40.75 in total [2], less than one leaderboard attempt at maximum effort [8], which is roughly the price of finding out whether the model clears your own ceiling.

What to watch

  • Anthropic publishing the harness and per-task budget behind 52.6 percent, which would make the figure reproducible outside its lab.
  • A leaderboard run of Fable 5.1 through Claude Code at maximum effort, giving a like-for-like cost per attempt against Fable 5's $14,180 across 210.
  • The same five tasks re-run at a cap near $67, which would separate Fable 5's four budget exits from genuine capability failures.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories