Skip to content

Topic

Model pricing and inference economics

The economics of AI model inference, covering per-token pricing, throughput and latency benchmarks, and batch versus real-time pricing differences.

Current stories

invest1 publisher

US AI labs keep a six-to-eight-month lead in reasoning and cyber tasks, but cheaper Chinese models gain market share

Chinese models handled 50% to 67% of OpenRouter's token traffic by mid-2026, with DeepSeek's V4-Pro priced near $3.96 per million output tokens. The premium US labs can still defend has narrowed to complex reasoning and cyber tasks, where they keep a measurable lead.

Reality

Evidence35
Adoption50
Hype gap+25
Incentives
Insufficient
Confidence35
invest4 publishers

Sonnet 5.5's extra tokens shrink its half-price edge over Opus to roughly a fifth at max effort

Anthropic's Claude Sonnet 5.5 beats Opus 5.5 at coding for half the per-token price, on its own tests and on Artificial Analysis's. At max effort it writes 60% more tokens per task, so moving coding work down a tier saves nearer a fifth than a half.

Perspective Coverage

4 publishers
Builder
Builder 39%
Operator
Operator 36%
Investor
Investor 25%

Reality

Evidence60
Adoption30
Hype gap+25
Incentives65
Confidence58
invest13 publishers

Google prices its third-ranked Gemini 4 Argon at half the cost of Anthropic's Claude Opus 5.5

Google priced Gemini 4 Argon at half Claude Opus 5.5's per-token rate for a model one composite index ranks third, behind two Anthropic models. On price alone it only matches OpenAI's newly discounted GPT-6.1 Sol, so the price edge Google is selling is against Anthropic.

Perspective Coverage

13 publishers
Builder
Builder 37%
Operator
Operator 31%
Investor
Investor 32%

Reality

Evidence55
Adoption20
Hype gap+30
Incentives65
Confidence55
invest5 publishers

OpenAI's near-$70 billion run rate barely clears the pace Anthropic hit in July

OpenAI's annualized revenue run rate is nearing $70 billion, up more than 70% since the quarter began, Axios reported. That is only about $5 billion above Anthropic's July rate, and Anthropic's figure has kept rising, so ranking the two labs has to wait for their prospectuses.

Perspective Coverage

5 publishers
Builder
Builder 15%
Operator
Operator 25%
Investor
Investor 60%

Reality

Evidence45
Adoption55
Hype gap+35
Incentives60
Confidence50
product5 publishers

OpenAI's 10-cent GPT-6.1 Sol price covers only cached input

OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output; the 10-cent figure in early coverage is its cached-input rate. Teams moving work off Astra should budget on the list rates and OpenAI's per-task costs.

Perspective Coverage

5 publishers
Builder
Builder 44%
Operator
Operator 34%
Investor
Investor 22%

Reality

Evidence50
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
invest9 publishers

Anthropic's reported $11.6B quarter puts OpenAI's 18% on the defensive before either lists

Reported quarterly revenue of $11.6 billion against OpenAI's $6.7 billion resets the enterprise-AI question. The pricing and concentration data underneath it flatter neither company.

Perspective Coverage

9 publishers
Builder
Builder 13%
Operator
Operator 21%
Investor
Investor 66%

Reality

Evidence55
Adoption60
Hype gap+30
Incentives75
Confidence60
invest3 publishers

Alphabet stock rises 0.6% the same day judge rejects forced sale of its ad exchange

A federal judge refused to order Google to sell its ad exchange, which takes a forced sale off the table for the ad stack, and the small bounce that followed says the argument over Alphabet has moved to the token price sheet.

Perspective Coverage

3 publishers
Builder
Builder 27%
Operator
Operator 25%
Investor
Investor 48%

Reality

Evidence60
Adoption30
Hype gap+25
Incentives55
Confidence60
build8 publishers

Gemini 3.8 Flash's introductory price doubles on December 31, 2026

Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 29%
Investor
Investor 18%

Reality

Evidence58
Adoption35
Hype gap+22
Incentives72
Confidence62
invest3 publishers

TypeSafe must meter 23.8 trillion input tokens to bill a million dollars for Jev

Jev charges $0.042 per million input tokens and nothing for output, so TypeSafe AI's revenue moves only with the state that agents pass in. Vercel, Cloudflare, LangChain and Langfuse listed it within a week.

Publishers:cryptobriefing.comindianexpress.comvercel.com

Perspective Coverage

3 publishers
Builder
Builder 53%
Operator
Operator 27%
Investor
Investor 20%

Reality

Evidence50
Adoption35
Hype gap+35
Incentives50
Confidence45
product5 publishers

Price cuts minutes apart send agent routing back to the spreadsheet

Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.

Perspective Coverage

5 publishers
Builder
Builder 38%
Operator
Operator 37%
Investor
Investor 25%

Reality

Evidence55
Adoption20
Hype gap+25
Incentives70
Confidence60
product8 publishers

OpenAI halves the API price of Sol and Luna against GPT-5.6's promotional rates

OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.

Perspective Coverage

8 publishers
Builder
Builder 36%
Operator
Operator 42%
Investor
Investor 22%

Reality

Evidence40
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
product7 publishers

Box measured Claude Opus 5.5 using a third of the tokens Opus 5 needed

Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.

Perspective Coverage

7 publishers
Builder
Builder 41%
Operator
Operator 38%
Investor
Investor 21%

Reality

Evidence45
Adoption30
Hype gap+35
Incentives60
Confidence60
invest16 publishers

Sol's 27-cent benchmark task undercuts Opus 5 by more than eleven times

Anthropic and OpenAI shipped cheaper model tiers minutes apart on Tuesday. OpenAI halved Sol's posted token prices, and the only per-task cost comparison between the two labs so far comes from OpenAI itself.

Perspective Coverage

16 publishers
Builder
Builder 29%
Operator
Operator 35%
Investor
Investor 36%

Reality

Evidence62
Adoption25
Hype gap+25
Incentives70
Confidence60

Earlier coverage

  1. Crude sells off for a fifth session as US and Iran hold 'very good' talks at UN

    Invest · September 23, 2026 · 1 publisher

  2. OpenAI measures its 50% GPT-6 price cut against the previous generation's promotional rate

    Science · September 22, 2026 · 2 publishers

  3. ByteDance releases next-generation AI model with price cut of more than 60%

    Leadership · September 16, 2026 · 1 publisher

  4. OpenAI's cost-per-task argument buys Luna room for ten failed tries before it loses on price

    Invest · September 8, 2026 · 1 publisher

  5. Token efficiency absorbs GPT-6 Astra's 2.5x price increase inside the coding harness

    Science · September 4, 2026 · 1 publisher

  6. OpenAI's top model at $4/$20 is a three-month answer to a permanent build decision

    Build · August 25, 2026 · 1 publisher

  7. Anthropic meters AWS Claude usage in one-cent units, and charges 10% more to pin inference to the US

    Build · August 18, 2026 · 1 publisher

  8. Developer habit, priced at $965B: what Anthropic's run actually proves

    Build · August 15, 2026 · 1 publisher