Published · 6h agoInvest7 min read
Anthropic cuts Opus 5.5 prices 20% on tokens and 60% on cached reads
Anthropic cut Opus 5.5 token prices 20% and cache reads 60% on September 22, with OpenAI answering within the hour. How much of either cut reaches a bill depends on a token mix only the buyer can measure.
Context for builders, not their beat.See today for builders

What happened
- Anthropic released Claude Opus 5.5 on September 22 and cut its price: input and output tokens are $4 and $20 per million, 20% below Opus 5, while cache reads drop 60% to $0.20 per million.
- Anthropic says Opus 5.5 performs at the level of Claude Fable 5.1 on most work and, combined with lower token use per task, costs 40% less to run than Opus 5 on typical workloads at default settings.
- Anthropic says the majority of agentic and coding work costs come from cached input.
- Simon Willison noted that Opus 4.5 through Opus 5 all shared the same $5/$25 pricing and that in longer agentic conversations more than 90% of input tokens are billed at cached rates.
- Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5, and is raising five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
Anthropic published two numbers on September 22 and they are not the same size. Input and output tokens fell 20 percent, to $4 and $20 per million [1]. Cache reads fell 60 percent, to $0.20 per million [1], which puts the old cache price at $0.50 [1]. Fresh input went from $5 to $4 [1][4], so a cached token now costs 5 percent of a fresh one where it used to cost 10 percent [2].
The second ratio is the one an agentic buyer pays. The independent developer Simon Willison noted that in longer agentic conversations more than 90 percent of input tokens are billed at cached rates [4], and Anthropic says the majority of agentic and coding work costs come from cached input [3]. Take that split: 0.9 times $0.20 plus 0.1 times $4 is $0.58 per million input tokens, against $0.95 under the old tariffs, a 39 percent cut on the input side alone [3]. Anthropic claims 40 percent lower cost to run at default settings, and attributes it to the lower per-token price plus fewer tokens used per task [2]. A buyer whose calls are short, one-shot and mostly fresh input collects the 20 percent and stops there [1].
OpenAI answered roughly an hour later with GPT-6 Sol and GPT-6 Luna at $2/$10 and $0.10/$0.50 per million, about half their GPT-5.6 equivalents by Willison's accounting [6]. Sol is exactly half Opus 5.5 on both sides of the sheet [4], and Opus 5.5's new $4/$20 is what GPT-5.6 Sol cost before that cut [7].
The 7x that is a 1.3x
The cleanest worked example of what a published cut is worth came a day earlier, from a team running a transactional email service. They swapped an admin classification call from a mid-size open model to TypeSafe AI's Jev and wrote it up on dev.to on September 21 [19]. Their cost figures were $0.52 per 1000 reviews against $0.07, a 7x cut [21]. Then they took the tariffs apart. Per input token the two providers are 1.31x apart, $0.055 against $0.042 per million; output accounts for 83 percent of the incumbent bill because output costs 15.5x input on that provider; and a footnote on Jev's usage dashboard says output is free [21]. The input share of the incumbent's bill, 17 percent of $0.52, is about $0.09 per 1000 reviews against Jev's $0.07 [11]. Of the quoted 7x, roughly 5.3x rests on that footnote [7]. The authors' own advice is to budget on the input price and treat free output as a discount that may not last [22].
Their speed number comes apart the same way. A median of 605ms against 7326ms is 12x [19]. But the incumbent writes a median of 459 output tokens per call against Jev's 138 [20], 3.3x more [9]. Divide latency by output tokens and the gap is roughly 3x, with two reasonable averaging methods landing 14 percent apart at 3.2x and 3.6x [20]. They quote the conservative figure, note that it still contains network, queueing and input read time, and say they have not tested whether shortening the prompt would recover most of the difference [20]. The team describes the 12x result as "fine and slightly boring" and says the more useful part of the exercise was that they produced two confident, wrong numbers first [26].
The per-task bill is a different number
Willison ran his recurring test, asking a model to generate an SVG of a pelican riding a bicycle at max thinking. Twice he got no response at all. Each run exhausted the model's 128,000 output token limit while still reasoning, at $2.56 and nearly 20 minutes a go [9]. That $2.56 is 128,000 tokens billed at $20 per million, so the pair cost $5.12 and returned nothing [5]. Willison wrote that the result makes him suspect max effort is effectively useless, and he is using Opus 5.5 as a default model in Claude Code anyway [9].
Artificial Analysis, which runs its own evaluations, scored the max-effort version 58 on its Intelligence Index, first of 212 models in its class, and also flagged it as verbose: 260M output tokens across the index against a median of 88M [8]. Priced at Opus 5.5's own $20 per million, that volume is $5,200 of output billing where a median-volume model bills $1,760 [6]. A 3x difference in tokens emitted swamps a 20 percent cut in the output tariff.
What the buyer actually controls
The dev.to team's real advantage is the switching cost they engineered. The model choice sits in a settings row rather than an environment variable, so reverting takes one request and no deploy [25]. The interactive admin button now calls Jev and returns in about 600ms with a verdict and a probability per category; background reviews, where nobody is waiting, still run on the incumbent [25].
Variance moved the interactive call: a standard deviation of 38ms for Jev against 2353ms for the incumbent, with a p95 of 11.5 seconds on a call an administrator waits on inside the request handler [23]. The sample has a hole in it, and they say so. Three of fifty calls to the incumbent returned text the JSON parser could not read, which on that route produces an error page, so the latency and cost figures cover only the 47 complete pairs, and the authors say they cannot tell which way that biases the numbers [24].
What the cheaper token comes with
Opus 5.5 is the first Opus model to launch with safeguards comparable to Claude Fable 5.1 on cybersecurity, biology and distillation. Most cybersecurity tasks are re-routed to Opus 4.8, and biology work impeded by those safeguards requires application to a new Life Sciences Verification Program [11]. A security team buying at $4/$20 is, on most of its tasks, buying Opus 4.8. Speed is a separate sheet again: fast mode costs $8 per million input and $40 per million output for up to 2.5x speed [17], which is 2x the price for 2.5x the throughput [8]. Anthropic says Opus 5.5 also generates output more than 30 percent faster than Opus 5 at the standard price, and is raising five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans [5][1]. On alignment, the company says the model scored better than any recent Claude model on its automated behavioral audit across nearly 2,000 scenarios. It also says it sees signs the model often suspects it is being evaluated, which limits how well pre-deployment testing predicts real-world behaviour [12].
The awkward part
Four days before that price cut, four paying subscribers sued Anthropic, OpenAI, SpaceXAI and Google in the US District Court for the Northern District of California. The suit alleges an "illegal" agreement to slow the pace of AI development, and argues that coordinated deceleration reduces the value consumers get for paid subscriptions, according to an Associated Press report carried by livemint.com [28][10]. The complaint locates the coordination on September 12, when Dario Amodei published an essay urging industry-wide cooperation on decelerating AI advances and Sam Altman, Elon Musk and Demis Hassabis each publicly responded in agreement the same day [29]. Anthropic's own product page opens with a line the plaintiffs will read closely: "Claude Opus 5.5 is our first release since we called for pacing the frontier." [15] Representatives for the four companies did not immediately respond to the AP's request for comment on Saturday [33]. Altman said on social media that OpenAI welcomes a "federal framework that sets consistent safety requirements," but added that "we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence" [32]. Amodei, in the essay itself, wrote that the government would need to "issue a narrow waiver for certain kinds of safety conversations" [31].
One: this is ordinary competition. A supplier that can serve a model on less compute passes some of it through within an hour of a rival's cut. Anthropic says that is what happened, stating that Opus 5.5 requires less compute to serve and that the pricing reflects that [16][6]. Two: some of the cutting is promotional in the same way Jev's free output is promotional, and the tariff to plan around is the one that survives the promotion [22]. Three: the plaintiffs are right about the September 12 exchange, in which case a price cut priced off cheaper serving cost is neutral evidence on whether capability is being held back [28][29].
I would plan on the second and price the first as upside. The falsifiable version: if the dev.to workload stays on Jev once output becomes billable, the engineering case was real; if it reverts, the 7x was a price sheet. A buyer who keeps a second supplier wired into a settings row holds an option, and the dev.to team put the cost of exercising it at one request and no deploy [25].
What to watch
- Whether Claude Sonnet 5.5 and Haiku 5.5 ship with the same $0.20-style cache read ratio or only the 20% headline cut.
- Whether the footnote on Jev's usage dashboard saying output is free survives, and what the dev.to team's $0.07 per 1000 reviews becomes if it does not.
- Whether the Northern District plaintiffs' damages theory holds up against a 20% token price cut and raised five-hour usage limits announced four days after filing.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Anthropic released Claude Opus 5.5 on September 22 and cut its price: input and output tokens are $4 and $20 per million, 20% below Opus 5, while cache reads drop 60% to $0.20 per million.
ReportedView cited source - [2]
Anthropic says Opus 5.5 performs at the level of Claude Fable 5.1 on most work and, combined with lower token use per task, costs 40% less to run than Opus 5 on typical workloads at default settings.
ReportedView cited source - [3]
Anthropic says the majority of agentic and coding work costs come from cached input.
ReportedView cited source - [4]
Simon Willison noted that Opus 4.5 through Opus 5 all shared the same $5/$25 pricing and that in longer agentic conversations more than 90% of input tokens are billed at cached rates.
ReportedView cited source - [5]
Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5, and is raising five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans.
ReportedView cited source - [6]
OpenAI released GPT-6 Sol and GPT-6 Luna roughly an hour after Anthropic's announcement, at $2/$10 and $0.10/$0.50 per million input and output tokens respectively, about half their GPT-5.6 equivalents by Willison's accounting.
ReportedView cited source - [7]
Opus 5.5's new $4/$20 matches what GPT-5.6 Sol cost before OpenAI's cut.
ReportedView cited source - [8]
Artificial Analysis scored the max-effort version of Opus 5.5 at 58 on its Intelligence Index, ranking it first of 212 models in its class, but also flagged it as verbose, generating 260M output tokens across the index against a median of 88M.
ReportedView cited source - [9]
Running his recurring test asking models to generate an SVG of a pelican riding a bicycle, Willison found Opus 5.5 at max thinking failed twice to return any response, exhausting the model's 128,000 output token limit while still reasoning; each failure cost him $2.56 and took nearly 20 minutes, he wrote that the result makes him suspect max effort is effectively useless, and he is nonetheless now using Opus 5.5 as a default model in Claude Code.
ReportedView cited source - [11]
Opus 5.5 is the first Opus model to launch with safeguards comparable to Fable 5.1 on cybersecurity, biology, and distillation, meaning most cybersecurity tasks are re-routed to Opus 4.8 and biology work impeded by those safeguards requires application to a new Life Sciences Verification Program.
ReportedView cited source - [12]
Anthropic says Opus 5.5 scored better than any recent Claude model on its automated behavioral audit across nearly 2,000 scenarios, while cautioning that it sees signs the model often suspects it is being evaluated, which limits how well pre-deployment testing predicts real-world behavior.
ReportedView cited source - [13]
Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks.
ReportedView cited source - [15]
Anthropic wrote: "Claude Opus 5.5 is our first release since we called for pacing the frontier."
ReportedView cited source - [16]
Anthropic says Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that.
ReportedView cited source - [17]
Fast mode for Opus 5.5 is available in Claude Code and the Claude Platform with up to 2.5x speed, and costs $8 per million input tokens and $40 per million output tokens.
ReportedView cited source - [19]
A side-by-side test of TypeSafe AI's Jev against a mid-size open model already in production returned a median latency of 605ms versus 7326ms, a 12x gap, according to a post published September 21 on dev.to by the DevOps Daily team, with measurements taken September 19, 2026 on the application host across 50 accounts at a transactional email service.
ReportedView cited source - [20]
The existing model writes a median of 459 output tokens per call against Jev's 138, and dividing latency by output tokens puts the gap at roughly 3x, with two reasonable averaging methods landing 14% apart at 3.2x and 3.6x; the authors quote the more conservative figure, caution that the adjusted number still contains network, queueing and input read time, and say they have not tested whether shortening the prompt would recover most of the difference.
ReportedView cited source - [21]
The dev.to team puts the existing model at $0.52 per 1000 reviews and Jev at $0.07, a 7x difference; per input token the two tariffs are 1.31x apart, at $0.055 and $0.042 per million; output accounts for 83% of the current bill because output costs 15.5x input on that provider, and a footnote on Jev's usage dashboard says output is free.
ReportedView cited source - [22]
The dev.to authors separate the engineering difference from the billing one and advise budgeting on the input price, treating free output as a discount that may not last.
ReportedView cited source - [23]
Variance rather than average speed was what the team says mattered: a standard deviation of 38ms for Jev against 2353ms for the existing model, on a call an administrator waits on inside the request handler, with a p95 of 11.5 seconds.
ReportedView cited source - [24]
Three of fifty calls to the existing model returned text the JSON parser could not read, which on that route produces an error page; those three calls recorded no timing, so the latency and cost figures cover only the 47 complete pairs, and the authors say they cannot tell which way that biases the numbers.
ReportedView cited source - [25]
The admin button now calls Jev and returns in about 600ms with a verdict and a probability for each category; background reviews, where nobody is waiting, still use the existing model, and the choice sits in a settings row rather than an environment variable so reverting takes one request and no deploy.
ReportedView cited source - [26]
The dev.to authors describe the 12x latency result as "fine and slightly boring" and say the more useful part of the exercise was that they produced two confident, wrong numbers first.
ReportedView cited source - [28]
A lawsuit filed Friday, September 18, in the US District Court for the Northern District of California by four named plaintiffs who each pay for a subscription to ChatGPT, Claude, Grok or Gemini accuses Anthropic, OpenAI, SpaceXAI and Google of making an "illegal" agreement to slow the pace of their AI development, arguing it violated antitrust laws and would reduce the value consumers get for paid AI subscriptions, according to an Associated Press report carried by livemint.com.
ReportedView cited source - [29]
The complaint argues the alleged coordination largely took place on September 12, when Anthropic CEO Dario Amodei published an essay urging industry-wide cooperation on decelerating AI advances in favour of stronger safety measures, and OpenAI CEO Sam Altman, SpaceXAI CEO Elon Musk and Google DeepMind co-founder and chair Demis Hassabis each publicly responded in agreement that same day.
ReportedView cited source - [31]
Amodei acknowledged potential antitrust problems in the essay, writing that the US government would need to "issue a narrow waiver for certain kinds of safety conversations".
ReportedView cited source - [32]
Altman said on social media that OpenAI welcomes a "federal framework that sets consistent safety requirements," but added that "we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence".
ReportedView cited source - [33]
Representatives for Anthropic, OpenAI, Google and SpaceXAI did not immediately respond to the Associated Press's request for comment on Saturday.
ReportedView cited source - [d1]
A cache read price of $0.20 per million that is 60% below the prior price implies the prior price was $0.50 per million.
Derived - [d2]
A cached input token now costs 5% of a fresh input token ($0.20 against $4), where it previously cost 10% ($0.50 against $5).
Derived - [d3]
At a 90% cached / 10% fresh input split, a million input tokens costs $0.58 under the new tariffs against $0.95 under the old ones, 39% lower.
Derived - [d4]
GPT-6 Sol's $2/$10 is exactly half Opus 5.5's $4/$20 on both input and output.
Derived - [d5]
Willison's $2.56 per failed run is 128,000 output tokens billed at $20 per million, and the two failures together cost $5.12 for no returned output.
Derived - [d6]
At $20 per million output tokens, 260M output tokens is $5,200 of output billing, against $1,760 for a model emitting the index median of 88M.
Derived - [d7]
With the two input tariffs only 1.31x apart, about 5.3x of the quoted 7x cost difference depends on Jev's free output.
Derived - [d8]
Opus 5.5 fast mode costs twice the standard token price on both input and output for up to 2.5x speed.
Derived - [d9]
The incumbent model's median output of 459 tokens per call is 3.3x Jev's 138.
Derived - [d10]
Anthropic released Opus 5.5 with its price cut four days after the antitrust complaint was filed.
Derived - [d11]
The input share of the incumbent's bill, 17% of $0.52 per 1000 reviews, is about $0.09, against Jev's all-in $0.07.
Derived
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- anthropic.comyesterdayIntroducing Claude Opus 5.5 \ Anthropic
- pivotnews.aiyesterdayDev.to team measures 12x speed gain from Jev, not 200x
- pivotnews.aiyesterdayAnthropic cuts Claude Opus 5.5 token prices 20% as OpenAI answers within the hour
- pivotnews.aiyesterdayAntitrust suit says AI leaders struck illegal deal to slow development