Product1 publisher2 min readPublished
Anthropic prices Opus 5.5 output tokens at twice OpenAI's rate for Sol
Anthropic trimmed Opus's token prices by a fifth and says its own benchmarks give it the agentic coding lead. OpenAI's GPT-6 Sol undercuts it on output tokens and claims to match Fable 5.1 on code.
The Product Desk · Product desk

What happened
- OpenAI put GPT-6 Sol at $2 for input and $10 for output, and GPT-6 Luna at $0.10 and $0.50, both trained with similar methods to GPT-6 Astra.
- Anthropic's own benchmarks put Opus 5.5 ahead of GPT-6 Astra at agentic coding on Terminal-Bench 4.0 and FrontierCode v1.1 (Main).
- Sol and Luna reach ChatGPT Work and Codex for Plus, Pro, Business and Enterprise customers, while free users and Go subscribers get only Luna in the desktop app.
- Each company also published its own proposed criteria for third-party evaluators judging the progress and safety of AI development.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- cost A team sending agent output to Opus 5.5 pays twice Sol's rate and 40 times Luna's for the same token volume. Price now decides which model gets the agent traffic as much as quality does.
- decision The comparison that would decide which model gets the agent traffic was run by Anthropic on Anthropic's model. Any team moving traffic has to reproduce it on its own repositories first.
- constraint A rollout designed around Sol's quality cannot reach free or Go seats at all. Access has to be gated by subscription tier before anyone writes the internal guidance.
- precedent Both labs cut list prices on iterative models in the middle of the pacing-the-frontier conversation, and procurement can now reasonably plan on the next iteration being cheaper again.
Somebody will open a billing dashboard on Monday and file a ticket to change one model string, because Opus 5.5 lists 20 percent under Opus 5 on both input and output [1]. That swap holds up only if the new model spends about the same number of tokens on the same job, and Engadget's account does not include a token count per task. A 20 percent per-token cut disappears when output volume rises 25 percent, since 1 divided by 0.8 is 1.25 [3].
The number a finance team can check is cost per completed task, and token volume is half of it. Anthropic's pitch for Opus 5.5 is complex work, "finding and fixing inefficiencies in software" and financial analysis and business work [2]. Those are long agent runs, and long runs are where output volume moves. The buyer this changes is the one running agents at volume on a cloud bill somebody else audits.
OpenAI's comparison is self-run too. The company says GPT-6 Sol matches Fable 5.1's coding performance at a lower cost, and that Sol makes about half as many factual mistakes as its predecessor [8][7].
The behavioural claims arrive the same way. Anthropic says Opus 5.5 "attempted to circumvent boundaries around 85 percent less often" than past models [4]. OpenAI grades its GPT-6 models on its own tests and says they lie less about the results of their coding work and refuse more attempts to bypass safeguards [10].
Some of the efficiency is plumbing. OpenAI says prompt caching improvements let Sol and Luna reuse more context and respond faster [9]. Opus 5.5 reaches developers through Claude, Amazon Web Services, Google Cloud and Microsoft Azure [12]. And OpenAI's discount is measured against the promotional pricing of the GPT-5.6 counterparts [6].
A test that settles the routing question: take ten tickets the agent already closed last month, re-run them on each candidate, and record output tokens, time to a merged fix, and how many runs needed a person to step in. Price per token is the only one of those numbers either company published [1][5]. Move the traffic only where output volume stays within 25 percent of the old model and the fix rate holds [3].
What to watch
- An independent run of Terminal-Bench 4.0 and FrontierCode v1.1 (Main) by anyone other than Anthropic.
- What Sol and Luna cost once the GPT-5.6 promotional pricing they are discounted against comes off the page.
- Whether any third-party evaluator adopts the criteria the two companies proposed, or publishes its own.