LeadershipNot yet confirmed elsewhere1 publisher3 min readPublished
Anthropic ships a price dial with its new model, and that is now the buying decision
Claude Opus 5 is pitched as near-frontier at half the cost. The cost multiple moves with the workload, and every figure on offer is the vendor's own.
The Board Room · Leadership desk
What happened
- Anthropic released Claude Opus 5, positioned as close to the frontier intelligence of its Fable 5 model at half the price.
- The model carries an effort setting customers use to trade intelligence against token consumption and speed.
- The company also states Opus 5 still trails Mythos 5 on cybersecurity tasks.
- Opus 5 becomes the default model on Claude Max and the strongest option on Claude Pro.
Why it matters
- decision Model selection stops being an annual choice and becomes a recurring per-workflow assignment of effort tiers, which someone has to own and revisit.
- cost Because the gain arrives as capability at an unchanged price rather than a price cut, teams that leave defaults alone keep paying the old bill and finance sees no line-item saving.
- constraint With no independent numbers on the table, any tiering plan has to be justified on an internal harness, which is engineering time nobody budgeted for.
- precedent Rivals now have to answer on cost per completed task rather than top-line scores, which pushes the whole comparison onto ground buyers cannot audit from a chart.
The cost advantage does not travel between workloads. On CursorBench 3.2 at maximum effort, Anthropic reports Opus 5 landing within 0.5% of Fable 5's peak score at half the cost per task [4]. On OSWorld 2.0, a computer use benchmark, it says Opus 5 beats Fable 5's best result at just over a third of the cost [12]. Those two are a 2x and a roughly 3x saving on the same comparison [17]. A savings model built on the headline "half the price" [9] will therefore be wrong in both directions, and which direction depends on what your teams actually do all day.
The price side is flat. Anthropic says Opus 5 costs the same as Opus 4.8 [2] while more than doubling its predecessor's Frontier-Bench v0.1 score at a lower cost per task [3], which puts performance per dollar on that benchmark at better than 2x for an unchanged list price [16].
The tiers also do not stack. The near-parity result comes from max effort on one benchmark [4], while the claim that even the cheapest setting beats every rival comes from Zapier AutomationBench, where Anthropic also reports a pass rate around 1.5x the next-best model at the same cost per task [6]. Those are different dials on different tests [19]. A team that reads the parity number and then runs its whole estate at low effort has bought neither result.
The verification problem sits underneath all of it. Every figure here, including the FrontierCode 1.1 comparison reported through an early-access customer [8], comes from Anthropic's own announcement [18]. The vendor is also the one telling you where it loses: it concedes cybersecurity tasks to Mythos 5 [10] while claiming state of the art on coding and knowledge work evaluations [1]. That is a useful admission, because it means "best model" is scoped to a task class, and making Opus 5 the standing default on Claude Max [14] carries an exception the supplier named before any buyer did.
The qualitative material is more interesting than the charts for tier selection. Anthropic describes a Frontier-Bench task where the model, denied any way to view a drawing, wrote its own computer vision pipeline to extract the geometry, and says no competing model in the same setup solved it in five attempts [15]. Persistence of that kind is exactly what a low effort setting is meant to ration. The research gains follow the same pattern of narrow, checkable deltas: 10.2 percentage points over Opus 4.8 on inferring molecular structures from spectroscopy data, 7.7 on predicting the functional effect of protein sequence variation [7].
So the number that matters is not on the vendor's chart. It is your own cost per completed task at each effort setting, on your own work, and nobody outside your organisation can compute it.
What to watch
- Whether any independent evaluator reproduces the CursorBench 3.2 and OSWorld 2.0 cost-per-task claims outside Anthropic's own harness.
- Whether competing vendors respond with list-price cuts or with their own effort-tier controls.
- Whether effort-tier selection starts appearing in enterprise contracts and published price sheets rather than only in API parameters.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence30
- Adoption32
- Hype gap+34
- Incentives88
- Confidence52
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Anthropic says that on coding and knowledge work evaluations such as Frontier-Bench and GDPval-AA, Opus 5 is the new state of the art.
ReportedSupportedSource: Anthropic announcement2 sources— create a free account to open themView cited source - [2]
Anthropic says Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8.
ReportedSupportedSource: Anthropic announcement2 sources— create a free account to open themView cited source - [3]
On Frontier-Bench v0.1, Anthropic says Opus 5 surpasses all other models and more than doubles Opus 4.8's performance at a lower cost per task.
ReportedSupportedSource: Anthropic announcement2 sources— create a free account to open themView cited source - [4]
On CursorBench 3.2 at max effort, Anthropic says Opus 5 performs within 0.5% of Fable 5's peak score at half the cost per task.
ReportedSupportedSource: Anthropic announcement2 sources— create a free account to open themView cited source - [5]
Anthropic says Opus 5 achieves greater performance at a given cost than all other models on the high, xhigh and max effort settings of CursorBench 3.2.
ReportedSupportedSource: Anthropic announcement2 sources— create a free account to open themView cited source - [6]
On Zapier AutomationBench, Anthropic says Opus 5's pass rate is around 1.5 times the next-best model for the same cost per task, and that even at its lowest effort setting it passes more tasks than any other model.
ReportedSupportedSource: Anthropic announcement2 sources— create a free account to open themView cited source - [7]
Anthropic says Opus 5 scores 10.2 percentage points higher than Opus 4.8 on its internal benchmark for inferring molecular structures from spectroscopy data, and 7.7 percentage points higher on predicting how variations in a protein's sequence affect its function.
ReportedSupportedSource: Anthropic announcement2 sources— create a free account to open themView cited source - [8]
An early-access customer report in the announcement says that on FrontierCode 1.1 Opus 5 approaches Fable-level performance at half the cost, and that within Devin it shows particular strength on difficult debugging and root-cause analysis tasks.
ReportedSupportedSource: Early-access customer quoted by Anthropic2 sources— create a free account to open themView cited source - [9]
Anthropic says Claude Opus 5 is available today and comes close to the frontier intelligence of Claude Fable 5 at half the price.
- [10]
Anthropic says Opus 5 remains behind Mythos 5 on cybersecurity tasks.
- [11]
Opus 5 has an effort setting that customers can use to optimise for intelligence or to conserve tokens for faster and cheaper results.
- [12]
On OSWorld 2.0, a computer use benchmark, Anthropic says Opus 5 outperforms every other model at any given cost and surpasses Fable 5's best result at just over a third of the cost.
- [13]
On ARC-AGI 3, Anthropic says Opus 5's score is three times as high as the next-best model.
- [14]
Opus 5 is the new default model on Claude Max and the strongest model on Claude Pro.
- [15]
Anthropic describes a Frontier-Bench task in which Opus 5, given a drawing of a machine part but no way to view it directly, wrote its own computer vision pipeline to pull the geometry from raw pixels and reconstructed the part, and says no competing model with the same setup solved it after five attempts.
- [16]
Because list price is unchanged against Opus 4.8 while Frontier-Bench v0.1 performance more than doubles at a lower cost per task, performance per dollar on that benchmark improves by more than 2x with no price cut.
- [17]
The reported cost advantage over Fable 5 differs by workload: half the cost per task on CursorBench 3.2 (a 2x ratio) versus just over a third of the cost on OSWorld 2.0 (roughly a 3x ratio).
- [18]
All benchmark and customer figures available for this release come from Anthropic's own announcement; no independent test result was supplied.
- [19]
The headline comparisons come from different effort settings on different benchmarks: max effort on CursorBench 3.2 for near-parity with Fable 5, lowest effort on Zapier AutomationBench for beating all rivals, so the two results cannot be assumed to hold together.
Sources
1 independent publisher whose own reporting we read for this story.
- anthropic.comIntroducing Claude Opus 5 \ Anthropic
1 article · August 26, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Vendor Benchmark TransparencyFollow
- Computer-use agentsFollow
- Frontier Model ReleasesFollow
- Agentic Coding and Long-Horizon TasksFollow
- Inference Effort and Token Budget ControlsFollow
- Model Price-Performance CompetitionFollow
Entities
- AnthropicFollow
- Claude Opus 5Follow
- Claude Opus 4.8Follow
- Claude Fable 5Follow
- Mythos 5Follow
- Claude Opus 4.7Follow
- Frontier-BenchFollow
- GDPval-AAFollow
- CursorBenchFollow
- OSWorld 2.0Follow
- Zapier AutomationBenchFollow
- ARC-AGI-3Follow
- FrontierCode 1.1Follow
- DevinFollow
- CursorFollow
- ZapierFollow
- LovableFollow
- FreeCADFollow
- Claude MaxFollow
- Claude ProFollow