Invest1 distinct publisher3 min readPublished
The top four slots on Artificial Analysis are the advertisement. The line item Anthropic actually moved is the one that scales with how long an agent runs, and its own savings estimate backs out that share at about 60 percent.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Divide the saving by the discount and the shape of an agent's bill falls out. If cutting the cache-read price by 75 percent takes 45 percent off a heavily agentic workload, then cache reads were around 60 percent of what that workload cost the week before [1]; on the average workload, where Anthropic's estimate is 25 percent, about a third [2]. Nothing else on the price list moved, with input holding at $10 per million and output at $50, twice Opus 5's rates and five times Sonnet 5's [9][4]. What changed is the ratio inside the invoice, because a cached token now prices at a fortieth of a fresh input token rather than a tenth [3], which is a reasonably precise statement about which workload Anthropic wants to win.
The complaint being answered is a procurement complaint. Cryptopolitan's account has Fable 5 adoption lagging because companies balked at bills they could not predict [12], and the unpredictable part of an agent's bill is the re-reading, the same repository and the same instruction block pulled back through the context window every time the loop comes round [11]. This is probably wrong, but I read the cut as cheap for Anthropic rather than generous, since the customer who was already going to run a two-hour agent now runs a three-hour one and the run rate arrives back where it started. The counter-thesis belongs in the same sentence: variance is a volume problem rather than a unit-price problem, so if spend per customer climbs while the per-workload saving holds at 25 percent, Anthropic will have sold a discount and collected a raise [11].
Which is why the capability numbers read as pricing support. Millennium told Anthropic the model traced a rare crash to a bug that had defeated its engineers for four or five years [15], and Browserbase put completion on its hardest browser-agent test at 82 percent against 74 for Opus 5 [16], a move that takes the failure rate from 26 to 18 and so removes 31 percent of failed tasks [5]. A two-hour run that fails is tokens billed against nothing shipped, so a failure rate is a cost line before it is a capability claim, and that is the defensible case for paying double Opus 5 [4]. The research and coding figures behind all this are Anthropic's own [7][8].
Anthropic is holding back the full capability from general release. Mythos 5.1 is the same underlying model behind separate safeguards, and it reaches only vetted cybersecurity and life-sciences groups through Project Glasswing [14], a channel that follows the July 23 pause on external cybersecurity evaluations after Claude got to real systems in tests meant to be sandboxed [13]. The thesis breaks if the write side of the cache, which our material does not price, absorbs the read-side saving, or if production hit rates land well under the ones Anthropic modelled [10][11]. A finance team can plan against that 60 percent figure, though Anthropic never published it itself [1].
Ranked by verification strength, evidence, and original report placement.
Anthropic released Claude Fable 5.1 on September 1, 2026, positioning it above its Haiku, Sonnet and Opus tiers for coding, knowledge work and long-running agentic tasks.
Fable 5.1 in 'max with fallback' mode took the top spot on Artificial Analysis's intelligence index with a score of 66, dethroning Claude Opus 5, while the 'xhigh with fallback' configuration scored 65.
Claude Opus 5 scores 63 on the Artificial Analysis index in both its max and xhigh modes.
Artificial Analysis is an independent party that ranks over 250 language models on price, speed and intelligence.
Anthropic holds the top four spots on the Artificial Analysis leaderboard, which also carries models from OpenAI, Google, SpaceXAI, Alibaba and DeepSeek.
The closest non-Anthropic entries are OpenAI's GPT-5.6 Sol in max mode and SpaceXAI's Grok 4.6, both scoring 61.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Fable 5.1's 52.6% science score arrives on a benchmark that was five days old1 distinct publisher
build
Anthropic cuts Fable 5.1 cache reads to a fortieth of its input price6 distinct publishers
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet relaying one vendor
Strip out the Artificial Analysis placement and every number in this story originates with Anthropic, restated by a single crypto news desk on launch day. Cryptopolitan does mark the Terminal-Bench figures as vendor-reported, which is more discipline than most launch write-ups manage — and also an admission that nobody re-ran them. The prices are the sturdy exception: they sit on a published rate card, which is why the cache-read arithmetic carries more weight here than the capability tables it is wrapped in.
Two hand-picked customer stories
Two users speak and Anthropic chose both: Millennium's long-unsolved bug and Browserbase's 82%. Neither harness is published, and neither figure is a usage number — no token volumes, no seat counts, no named migration off another model. The one claim about real-world uptake runs the other way, that Fable 5's adoption lagged on unpredictable bills, and it arrives with nothing attached. What is genuinely observable is a launch, a rate-card change and a gated second tier.
Ranking loud, arithmetic quiet
Three points on a composite index, earned in a 'max with fallback' configuration whose running cost nobody prices, does most of the headline work — and a rerun of the index could erase it. Beneath that sits the part that actually holds: cut one line item by 75%, watch a heavy agent's bill fall 45%, and that line item was about 60% of the bill. Anthropic effectively conceded the shape of its own cost structure and got covered for the leaderboard instead.
Vendor wrote every number
Anthropic supplied the benchmark table, selected which customers got quoted, and sized the savings estimate that makes its own price change sound generous — three layers of self-interest before a reader reaches a figure. The 25% and 45% claims do useful commercial work: they reframe a model still charging $10 and $50 as effectively cheaper than the Opus tier it sits above. Cryptopolitan's stake is smaller and visible, a newsletter pitch dropped mid-article.
Firm on price, provisional on rank
We would bet on the rate card and the arithmetic drawn from it. We would not bet on the ordering: one publisher, one launch day, no replication of the Terminal-Bench tables or Browserbase's harness means the capability story could look different at the next index refresh while $0.25 per million cached tokens stays exactly where it is. Read the economics as settled and the crown as contingent.