Leadership1 publisher3 min readPublished
The AI bill nobody reconciles: cost per finished task, not per million tokens
A new cost analysis puts OpenAI's frontier model at half Anthropic's price per benchmark task. The retry and cleanup arithmetic behind that number is less settled than the price sheet.
The Board Room · Leadership desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A doit.com analysis argues cost per token is the wrong unit for agentic and production workloads, and that the right unit is expected cost per completed task plus the cost of cleaning up wrong outputs that escape automated checks.
- On raw per-task token cost in mid-2026, Artificial Analysis measured GPT-5.6 Sol at $1.04 per Intelligence Index task versus Claude Opus 5 at $2.03 and Claude Fable 5 at $2.75.
- Flagship list prices have converged: Claude Opus 5 costs $5 input / $25 output per million tokens and GPT-5.6 Sol costs $5 / $30, making Anthropic cheaper on output list price and rendering the older "Claude costs more per token" framing largely obsolete at the frontier.
- Anthropic's tokenizer from Opus 4.7 onward produces roughly 30 percent more tokens for the same text, and Opus 4.8 and Opus 5 report about 1.88 tokens per English word versus about 1.17 on GPT-5's o200k encoding, so a given dollar-per-million rate buys fewer words of Claude context.
- Reasoning tokens are billed as output tokens on both platforms, OpenAI hides them while Anthropic can return them, and a 500-token visible answer can carry thousands of billed reasoning tokens.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
A cost analysis published by doit.com argues that the unit of AI procurement has moved from dollars per million tokens to expected cost per completed task, plus the cost of cleaning up wrong outputs that escape automated checks [1]. That matters because on the old unit the cheaper frontier vendor has flipped, and on the new one the ranking depends on numbers most buyers do not currently measure [2][6].
Start with the price sheet, since that is where most comparisons stop. Anthropic's Claude Opus 5 lists at $5 per million input tokens and $25 output; OpenAI's GPT-5.6 Sol lists at $5 and $30, which makes Anthropic the cheaper option on output list price [3]. The framing that Claude costs more per token is, at the frontier, obsolete [3].
It is also irrelevant, because tokens are not a common currency. Anthropic's tokenizer from Opus 4.7 onward produces roughly 30 percent more tokens for the same text, and Opus 4.8 and Opus 5 report about 1.88 tokens per English word against about 1.17 on GPT-5's o200k encoding [4]. Normalise the list prices to words and the advantage reverses: about $47 per million words of Claude output versus $35.10 for GPT-5.6 Sol, and $9.40 versus $5.85 on the input side [17][18]. Reasoning tokens are billed as output on both platforms, with OpenAI hiding them and Anthropic able to return them, so a 500-token visible answer can carry thousands of billed tokens [5].
Measured per task rather than per token, Artificial Analysis put GPT-5.6 Sol at $1.04 per Intelligence Index task in mid-2026, against $2.03 for Claude Opus 5 and $2.75 for Claude Fable 5 [2]. That is a 1.95x gap [14]. The reason the gap does not track list prices is that agentic work is a loop, not a call: Vantage models a representative 50-turn coding session at roughly 1 million input tokens and 40,000 output, a 25-to-1 ratio in which input is about 85 percent of the bill because the model re-reads an accumulating context every turn [7]. Gartner's March 2026 analysis found agentic workflows burn 5 to 30 times more tokens per task than a chatbot query [8]. A preprint by Bai et al. covering eight frontier models on SWE-bench Verified found agentic coding consumes up to 1,000 times more tokens than code chat, with up to 30x run-to-run variance [10]. Priced at list, that 50-turn session costs about $6.00 on Opus 5 and $6.20 on GPT-5.6 Sol [19].
The case for the pricier model is reliability. On tau-bench, Claude Opus 4.8 held far more of its single-attempt score across repeated runs and violated policy less than half as often as GPT-5.5 [6]. Since a failed trajectory still bills in full [13], and expected cost per completed task is attempt cost divided by success probability [16], a 1.95x cost premium needs roughly 1.95x the success rate to break even on retries alone [15]. The rest has to come from cleanup avoided, which is exactly the line item nobody instruments. Gartner's Will Sommer put the general risk this way: product chiefs "should not confuse the deflation of commodity tokens with the democratization of frontier reasoning" [9].
Watch whether vendors start publishing repeated-run reliability alongside price, and whether your own finance function can produce a retry count. Bai et al. note that the models themselves underestimate their token usage [11]; prompt cache reads cost 10 percent of input on both platforms, which changes long-session arithmetic more than any list-price move [12].