Leadership1 distinct publisher3 min readPublished
Claude Opus 5 is pitched as near-frontier at half the cost. The cost multiple moves with the workload, and every figure on offer is the vendor's own.
The Board Room · Leadership desk
Compiled by The Board RoomSomething wrong?How this is made
The cost advantage does not travel between workloads. On CursorBench 3.2 at maximum effort, Anthropic reports Opus 5 landing within 0.5% of Fable 5's peak score at half the cost per task [7]. On OSWorld 2.0, a computer use benchmark, it says Opus 5 beats Fable 5's best result at just over a third of the cost [9]. Those two are a 2x and a roughly 3x saving on the same comparison [16]. A savings model built on the headline "half the price" [1] will therefore be wrong in both directions, and which direction depends on what your teams actually do all day.
The price side is flat. Anthropic says Opus 5 costs the same as Opus 4.8 [5] while more than doubling its predecessor's Frontier-Bench v0.1 score at a lower cost per task [6], which puts performance per dollar on that benchmark at better than 2x for an unchanged list price [17].
The tiers also do not stack. The near-parity result comes from max effort on one benchmark [7], while the claim that even the cheapest setting beats every rival comes from Zapier AutomationBench, where Anthropic also reports a pass rate around 1.5x the next-best model at the same cost per task [10]. Those are different dials on different tests [19]. A team that reads the parity number and then runs its whole estate at low effort has bought neither result.
The verification problem sits underneath all of it. Every figure here, including the FrontierCode 1.1 comparison reported through an early-access customer [15], comes from Anthropic's own announcement [18]. The vendor is also the one telling you where it loses: it concedes cybersecurity tasks to Mythos 5 [3] while claiming state of the art on coding and knowledge work evaluations [2]. That is a useful admission, because it means "best model" is scoped to a task class, and making Opus 5 the standing default on Claude Max [12] carries an exception the supplier named before any buyer did.
The qualitative material is more interesting than the charts for tier selection. Anthropic describes a Frontier-Bench task where the model, denied any way to view a drawing, wrote its own computer vision pipeline to extract the geometry, and says no competing model in the same setup solved it in five attempts [14]. Persistence of that kind is exactly what a low effort setting is meant to ration. The research gains follow the same pattern of narrow, checkable deltas: 10.2 percentage points over Opus 4.8 on inferring molecular structures from spectroscopy data, 7.7 on predicting the functional effect of protein sequence variation [13].
So the number that matters is not on the vendor's chart. It is your own cost per completed task at each effort setting, on your own work, and nobody outside your organisation can compute it.
Ranked by verification strength, evidence, and original report placement.
Anthropic says that on coding and knowledge work evaluations such as Frontier-Bench and GDPval-AA, Opus 5 is the new state of the art.
Anthropic says Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8.
On Frontier-Bench v0.1, Anthropic says Opus 5 surpasses all other models and more than doubles Opus 4.8's performance at a lower cost per task.
On CursorBench 3.2 at max effort, Anthropic says Opus 5 performs within 0.5% of Fable 5's peak score at half the cost per task.
Anthropic says Opus 5 achieves greater performance at a given cost than all other models on the high, xhigh and max effort settings of CursorBench 3.2.
On Zapier AutomationBench, Anthropic says Opus 5's pass rate is around 1.5 times the next-best model for the same cost per task, and that even at its lowest effort setting it passes more tasks than any other model.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Vendor-only, partly unquantified
The cluster contains exactly one source: the vendor's launch post. It supplies specific, checkable-sounding figures (within 0.5% of Fable 5 on CursorBench 3.2, ~1.5x Zapier pass rate, 3x ARC-AGI 3, +10.2 and +7.7 percentage points on internal life-sciences benchmarks), which lifts it above bare assertion. But methodology, rival configurations, effort settings for baselines and absolute prices are absent, several results are only described as charts, and no independent reproduction exists in the supplied material.
Launch-day availability and partner pilots
Real adoption facts exist but are all launch-day and vendor-disclosed: general availability, promotion to default model on Claude Max and strongest model on Claude Pro, and named early-access partners (Devin/Cognition, Cursor, Zapier, Lovable, a genomics team, a trading firm) describing hands-on use. There is no independent usage, traffic, revenue or retention data, and pre-release pilot access is not production adoption.
Headline overstates a workload-dependent result
The 'near-frontier at half the price' headline is presented as one property of the model, but the cost multiple against Fable 5 varies by workload (roughly 2x on CursorBench 3.2, roughly 3x on OSWorld 2.0) and the flagship comparisons are drawn from different effort settings that cannot be assumed to hold together. Superlatives such as new state of the art and 3x ARC-AGI 3 are unaccompanied by methodology, and the sole disclosed weakness gets one clause. The gap is moderate rather than severe because the release is real, priced-in-line with its predecessor, and the numbers offered are specific.
Sole source is the seller's launch post
The only source is Anthropic's product announcement for a model it sells, including the choice of benchmarks, effort settings and comparison baselines. The supporting testimonials come from early-access commercial partners (Cursor, Zapier, Cognition/Devin, Lovable) whose products depend on the model and whose quotes are curated by the vendor. There is no adversarial or independent voice in the cluster.
Clear provenance, single perspective
Confidence in this assessment is moderate: the provenance of every claim is unambiguous and the vendor's own text is internally consistent, so the structural findings (vendor-only evidence, workload-dependent cost multiple, mixed effort settings) are solid. But with one publisher and no independent measurement, the accuracy of the underlying performance and cost claims cannot be judged, which caps confidence near the midpoint.
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
product
Anthropic's protein binders got tested by outside labs. The benchmark is still Anthropic's.1 distinct publisher
build
Gartner: agent inference cost rises 5x by 2028, so budget per workflow, not per model1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026