Build1 distinct publisher3 min readPublished
General capability barely moves in the rest of Anthropic's table, so the thing teams have to configure is a five-level effort dial that ran one SVG prompt from ten cents to $3.30. The per-token rate never changes.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The bottom of the effort dial does not behave like a dial. At `low`, Simon Willison's pelican prompt produced 1,998 output tokens in 23.8 seconds for 10.017 cents, with no summarized reasoning in the transcript [9]. At `medium` it produced 1,977 tokens in 23 seconds for 9.912 cents, 21 tokens fewer than `low` [10]. Claude's output token count includes reasoning tokens [11], so Willison reads both runs as skipping reasoning entirely on this prompt [12]. `high` bought a short planning summary for 2,612 tokens, 29.6 seconds and 13.087 cents [13]. The steps that cost real money are `xhigh`, at 36,767 tokens, 7 minutes 51 seconds and $1.83 [14], and `max`, at 65,927 tokens, 13 minutes 54 seconds and $3.30 [15].
Divide cost by output tokens at each level and the rate is flat, around $50 per million at `low`, `high`, `xhigh` and `max` [1]. The prompt is one sentence, so input tokens are noise inside those totals. Divide tokens by wall clock and throughput is flat too, between 78 and 84 tokens per second [8]. Effort turns out to be a token budget on one meter, not a price tier and not a different serving path. `max` costs 33 times `low` because it writes 33 times as many tokens [2][3], and it takes 35 times as long [4]. In my context the useful position is `high`, because it is the last one before latency crosses from seconds into minutes [13][14].
The 52.6% is a much bigger jump than anything in those runs: 27.9 points over Fable 5 and 30.2 points over GPT-5.6 Sol on the same board [5]. It is also a score on a 0.1-versioned benchmark first announced on 27 August, five days before the model shipped [5][6], and Willison's reading of the rest of the announcement is that other benchmarks improved only slightly [6]. For that number to transfer, your tasks need to resemble the terminal-driven science tasks in that harness, the harness needs to score them the way you would score them, and you need to call the model at whatever effort level produced the score. The writeup does not name that level, and these runs show effort moving output volume by a factor of 33 on this same model [3].
Willison is explicit that he lost confidence in the pelican prompt as a proxy for general model quality back in July, and now uses it mainly for comparisons within a model family and for one prompt across reasoning levels [16]. That is the honest use of it, and it is the same discipline a five-day-old benchmark has not had time to earn. `max` gave him the best pelican he has seen from any Anthropic model [15], still with less flair than Gemini 3.7 Flash [17]. $3.30 for an SVG is a rounding error right up until it sits inside a retry loop.
Ranked by verification strength, evidence, and original report placement.
Claude Fable 5.1 and Mythos 5.1 were released, written up on 1 September 2026.
The Anthropic announcement spends a notable amount of time on scientific research, boasting of a 52.6% score for Fable 5.1 on the brand new Terminal-Bench-Science 0.1 benchmark.
On Terminal-Bench-Science 0.1 the reported comparison scores are 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol.
At effort xhigh the run produced 36,767 output tokens in 7 minutes 51 seconds and cost $1.83.
At effort max the run produced 65,927 output tokens in 13 minutes 54 seconds and cost $3.30, and gave Willison the best pelican he has seen from any Anthropic model.
Anthropic say Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks".
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Anthropic cuts Fable 5.1 cache reads to a fortieth of its input price6 distinct publishers
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
leadership
Anthropic cuts Fable 5.1 prices by 25% and launches two-tier safeguard system with Mythos 5.11 distinct publisher
product
Ramp's July card data puts Opus 4.8 at 3.5 times Claude Fable's spend share1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Metered runs, relayed scoreboard
Two very different grades of evidence share this page. The effort-level figures are first-hand and unusually exact — tokens, seconds and fractions of a cent per run — and they hold up arithmetically: divide any cost by its token count and you get the same $50 per million. The 52.6% does not have that quality; it is Anthropic's figure, quoted, on a benchmark nobody in this reporting has run themselves. Willison also concedes the limits of his own instrument before he uses it, which is more candour than the vendor table offers.
One prompt, one plugin, one tester
Everything observable amounts to a launch day, a vendor benchmark table and five runs of a single SVG prompt through one developer's command-line tool, plus a $1.37 follow-up. No team deployment, no production traffic, no second practitioner repeating the sweep. That is enough to characterise the effort dial's cost curve and nothing like enough to say anyone is running Fable 5.1 at scale.
Headline outruns the table
"Sets a new standard" is doing work that a five-day-old science benchmark is being asked to underwrite while, by Willison's account, the rest of the scoreboard shifts only slightly. Meanwhile the change a user will actually feel — a dial that can turn a ten-cent request into a $3.30, fourteen-minute one, with no off setting — is not what the launch is being sold on. The overstatement is in the framing, not in the numbers: the measured parts of this story are, if anything, undersold.
Vendor supplies the yardstick
Anthropic is both the subject and the source of the number it wants read: a brand-new science benchmark on which its new model roughly doubles the field, published while the older benchmarks move a little. The reviewer's position is milder but not neutral — Willison maintains the llm-anthropic plugin he ran the tests through and patched it for this post, and the pelican test is his own invention, which he flags as a weakening signal rather than defending. No sponsorship or access arrangement is disclosed either way.
Exact figures, single witness
We can be fairly confident about what happened on one laptop on 1 September: the run-level numbers are precise, internally consistent, and the two copies we hold agree on every figure — the later revision changes only a judgement about backwards-spinning wheels. What we cannot stand behind is generality. One prompt, one attempt per level, one publisher, and a vendor benchmark taken on trust. Read the cost curve as solid and the capability story as provisional.