Build1 distinct publisher3 min readPublished
Anthropic charges the same for 5.1 as it does for Fable 5, so the upgrade itself is free. Whether the Terminal-Bench-Science jump from 24.7% to 52.6% reaches your queue depends on how much work Fable 5 was already failing.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Terminal-Bench-Science hands a model a shell and a set of multi-step research tasks, then scores the fraction it completes correctly [3]. A pass rate is fundamentally a property of the task distribution being sampled. The 27.9 points between the two scores [17] therefore sit entirely in tasks Fable 5 fails. On any task Fable 5 already finishes, the headroom is zero and 5.1 can only match it.
Had the four fixtures been drawn from the benchmark's own distribution, a model that completes 24.7% of it [2] would sweep all four about 0.4% of the time [18]. The fixtures instead sit outside that distribution, and that gap is the finding. The planted bugs sat behind a suite that failed loudly, and the bad experiment rows were described in lab notes the model was free to read [10][8]. Both fixtures are well specified. The Terminal-Bench-Science tasks Fable 5 misses are, by construction, the ones it does not finish.
Two conditions have to hold before that benchmark delta reaches a pipeline. The queue needs a real population of work Fable 5 currently fails, and each task needs enough repeats to show a rate rather than an outcome. One run per task cannot separate 24.7% from 52.6%, because both numbers predict passes and failures on the same task.
The measured differences are small and consistent. On the coding fixture 5.1 finished 3.6 seconds sooner, 21% off a 17.2-second run, at the same seven cents [11][19]. It saved 1.4 seconds and 1.4 cents on the research fixture, which comes to $14 per thousand runs [9][20]. The two reasoning problems cost it 425 fewer output tokens, 23% fewer, worth about two cents at $50 per million [12][21]. With pricing identical [4], none of that argues against switching. It argues against rebuilding anything to mark the occasion.
Dan Shipper of Every called 5.1 the strongest coding model his team has used, after a week of testing [5]. That holds alongside a four-task tie without much strain: a week of varied production work samples far more of the hard tail than four fixtures and a deliberately harder tiebreaker do [13], and the tail is where the 2.1x lives [16]. The New Stack's own caution stops short of an accusation. It notes that benchmark sets are scored under conditions vendors help define, and that tuning toward graded tests has happened, while declining to claim Anthropic did it here [15].
The weak point of the comparison is reproducibility: the prompts are useless without the folders of planted-error data behind them, so they were not published [14]. Take it as one workload, run once, across four tasks. That is still a better predictor of your own queue than a vendor table, and it is the shape of test worth copying, with fixtures built from work your current model gets wrong and enough repeats to report a percentage instead of an anecdote.
Ranked by verification strength, evidence, and original report placement.
The Fable 5.1 announcement highlights the Terminal-Bench-Science agentic research benchmark, where 5.1 scores 52.6% against Fable 5's 24.7%.
The sensor data audit was added as a tiebreaker after three rounds of perfect ties on accuracy, and was built to be harder than the first three tests; both models handled every trap, including shifting the fast clock back before filtering the time window.
Anthropic launched Claude Fable 5.1, calling it "our most advanced model for coding and knowledge work".
Terminal-Bench-Science gives a model a terminal and a set of multi-step scientific research tasks, then scores what percentage it completes correctly.
Fable 5.1 costs exactly what Fable 5 costs: $10 per million input tokens and $50 per million output tokens.
Every CEO Dan Shipper posted that after a week of testing, Fable 5.1 was "the strongest coding model we've used".
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 5, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Anthropic's Fable 5.1 moves the hard part from prompting to bounding what it may do14 distinct publishers
product
Anthropic bills Pro seats extra for the flagship model already in their picker1 distinct publisher
science
Token efficiency absorbs GPT-6 Astra's 2.5x price increase inside the coding harness1 distinct publisher
invest
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-hand runs, unrepeatable
Every number that matters comes from one reviewer's single pass at each task on his own machine, and they are precise: 19.2 seconds against 20.6, $0.086 against $0.100, 8 of 8 tests passing in 3 turns. The New Stack also says plainly that it cannot publish the prompts, because the tests only work with folders of data files carrying planted errors, so no reader can check the result. Anthropic's 52.6% and 24.7% are announcement figures, not numbers our coverage or anyone else has independently verified.
Launch-week signals only
What exists is a few days old: the model shipped at unchanged pricing, Dan Shipper reported a week of testing before release, Min Choi gathered first-day examples, and The New Stack ran its four tasks. The only observation with numbers attached is the one The New Stack made itself. Nothing in this reporting describes production traffic, migrations, or spend at any organisation.
Doubling upstream, dead heat downstream
The launch case is a science benchmark going from 24.7% to 52.6%. What the desk test produced is 24 of 24 for both models, with the visible gains measured in single-digit seconds and fractions of a cent. The most pointed number sits inside The New Stack's own copy rather than in anyone's framing: on the tiebreaker audit, built specifically to be hard, Fable 5 finished in 4 turns for $0.134 while 5.1 spent 5 turns and $0.304.
Vendor figure, vendor conditions
The one number showing a large difference originates in Anthropic's own launch post, and The New Stack names the mechanism rather than implying it: benchmarks cover narrow task sets under conditions vendors help define, and some companies have tuned models toward their graded tests, though the reviewer says he is not accusing Anthropic. The two amplifiers have audiences to serve — Shipper runs a subscription publication about AI workflows, Choi curates model demos. The reviewer's own pull goes the other way, since a tie makes for a flatter piece than a doubling.
Internally consistent, thin sample
The arithmetic holds together: token counts, turn counts and per-run costs line up with the timings and the posted per-million rates, and the reviewer flags his own limits, including a caveat about cache-read prices that the text cuts off mid-sentence. What keeps this mid-range is scale — one tester, one run per task, four tasks of his own design, and a vendor benchmark neither reproduced nor inspected.