Skip to content

Topic

LLM API Pricing

Pricing models providers use for LLM API access, including per-token rates, cached-input discounts, and peak/off-peak fee schedules.

Current stories

security9 publishers

Google gives vetted defenders a Gemini 4 Argon build without cyber guardrails

Google will give vetted defenders and its own teams a Gemini 4 Argon build with no cyber guardrails, saying the model finds and patches critical flaws unaided. Wiz is the first named outside user, and the bug-finding evidence published so far comes from Google's own internal tests.

Perspective Coverage

9 publishers
Builder
Builder 39%
Operator
Operator 39%
Investor
Investor 22%

Reality

Evidence50
Adoption25
Hype gap+35
Incentives75
Confidence60
build17 publishers

OpenAI courts builders with a cheaper model and ChatGPT's 1.2 billion weekly users

OpenAI priced GPT-6.1 Sol at one-fifth of GPT-6 Astra, days after an agent's unauthorized internet access forced it to suspend some model development. Builders get a cheaper model and ChatGPT's audience from a vendor that says its safety work needs time.

Perspective Coverage

17 publishers
Builder
Builder 48%
Operator
Operator 35%
Investor
Investor 17%

Reality

Evidence62
Adoption35
Hype gap+20
Incentives72
Confidence64
invest5 publishers

Cached context shrinks the discount from Anthropic's half-price Sonnet 5.5

Anthropic released Sonnet 5.5 at $2 and $10 per million input and output tokens, half the Opus 5.5 rate. How much a buyer saves by moving work down a tier depends on tokens burned per task and on cache reads priced identically on both models.

Perspective Coverage

5 publishers
Builder
Builder 46%
Operator
Operator 38%
Investor
Investor 16%

Reality

Evidence55
Adoption35
Hype gap+20
Incentives70
Confidence60
build1 publisher

Claude's new tokenizer and GPT-6's 272K price cliff break old LLM cost models

Claude 4.7 emits about 30 percent more tokens for the same text and GPT-6 bills roughly double above 272K input tokens, a dev.to digest reports. Budget checks built on old token counts now undercount, so prompt size needs a hard cap enforced in code.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence35
science4 publishers

Gemini 4 Argon costs 2.7 times as much per task as GPT-6.1 Sol at the same token price

Google's Gemini 4 Argon matches GPT-6.1 Sol's $2/$10 token price but costs 2.7 times as much per task, according to Artificial Analysis. Argon uses more tokens per job, so buyers still have to compare frontier models by cost per completed task.

Perspective Coverage

4 publishers
Builder
Builder 36%
Operator
Operator 34%
Investor
Investor 30%

Reality

Evidence68
Adoption15
Hype gap+20
Incentives55
Confidence65
product4 publishers

OpenAI's cheaper GPT-6.1 Sol leaves its new Dots agents on the pricier Astra model

OpenAI says GPT-6.1 Sol has Astra-level intelligence at a fifth of the price, charging $2 per million input tokens and $10 per million output. That saving reaches the API workloads teams build themselves, while the Dots agents launched with it run on Astra inside seat plans that got dearer.

Publishers:cnet.comlennysnewsletter.commashable.comstratechery.com

Perspective Coverage

4 publishers
Builder
Builder 35%
Operator
Operator 39%
Investor
Investor 26%

Reality

Evidence45
Adoption20
Hype gap+30
Incentives60
Confidence55
product11 publishers

Sonnet 5.5 comes within two points of Opus 5.5 at half the per-token price

Anthropic says Sonnet 5.5 nearly ties Opus 5.5, 1,844 to 1,846, on an everyday-work benchmark while running more than 30% faster than Sonnet 5. For teams paying double per token for Opus, Sonnet becomes the sensible default, with Opus kept for long, ambiguous jobs.

Perspective Coverage

11 publishers
Builder
Builder 36%
Operator
Operator 43%
Investor
Investor 21%

Reality

Evidence45
Adoption50
Hype gap+25
Incentives65
Confidence55
build8 publishers

Meta hands its Muse models and coding agent to an enterprise unit run by CJ Desai

Meta grouped Muse API, Muse Code and Business Agent into a new enterprise platform whose Muse Spark model costs developers $1.25 per million input tokens. The endpoint is ready to test today, while the enterprise package still has no published price or delivery date.

Perspective Coverage

8 publishers
Builder
Builder 26%
Operator
Operator 28%
Investor
Investor 46%

Reality

Evidence62
Adoption25
Hype gap+35
Incentives70
Confidence65
build17 publishers

Opus 5.5 matched Opus 5's puzzle answers for up to 69 percent less at its lower default effort

Opus 5.5 matched Opus 5 on two reasoning puzzles in The New Stack's tests at 43 to 69 percent lower cost. Both ran at default effort, medium on the new model and high on the old, so the saving a team sees depends on the effort level it pins.

Perspective Coverage

17 publishers
Builder
Builder 43%
Operator
Operator 33%
Investor
Investor 24%

Reality

Evidence62
Adoption48
Hype gap+22
Incentives58
Confidence58
build1 publisher

Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence50
build4 publishers

DeepSeek open-sources the harness, then raises the price of the model

Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.

Perspective Coverage

4 publishers
Builder
Builder 51%
Operator
Operator 31%
Investor
Investor 18%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence58

Earlier coverage

  1. OpenAI's August changelog cuts Sol prices and puts a date on them

    Build · August 27, 2026 · 2 publishers

  2. Opus 5.5 cost less per unit of coding work than Sonnet 5, with fewer review rounds needed

    Build · September 25, 2026 · 1 publisher

  3. Anthropic cuts Fable 5.1 prices by 25% and launches two-tier safeguard system with Mythos 5.1

    Leadership · September 1, 2026 · 3 publishers

  4. Google's Opus comparison for Gemini 3.8 Flash ran entirely inside its own coding tool

    Build · September 1, 2026 · 1 publisher

  5. Inception's 1,107 tokens per second needs a batch size before it enters your capacity plan

    Build · September 8, 2026 · 2 publishers

  6. Astra's Critical cyber rating ships a real-time pause switch inside the Bedrock service boundary

    Build · September 10, 2026 · 18 publishers

  7. Astra bills at long-context rates once a request passes 30 percent of its input window

    Build · September 11, 2026 · 1 publisher

  8. Default LLM calls pay list price for work that caching and batch would discount

    Build · September 25, 2026 · 1 publisher

  9. Opus 5.5 diverts most cybersecurity requests to the older Opus 4.8

    Leadership · September 22, 2026 · 2 publishers

  10. A tester left Claude Opus 5.5 running unattended for 18 hours across six repositories

    Security · September 23, 2026 · 3 publishers

  11. Sonnet's output rate caps a $20 seat at 1.3 million tokens a month

    Build · September 24, 2026 · 1 publisher

  12. Opus 5.5's claimed 40% cost cut needs a cache-heavy workload to appear

    Science · September 23, 2026 · 2 publishers

  13. OpenAI's 50 percent API price cut doubles the token volume a flat budget buys

    Security · September 23, 2026 · 1 publisher

  14. Anthropic cuts Opus 5.5 prices 20% on tokens, 60% on cache reads, citing fewer tokens burned for 40% total savings

    Invest · September 23, 2026 · 1 publisher

  15. OpenAI cuts prices on new GPT-6 Sol and Luna models

    Product · September 23, 2026 · 1 publisher

  16. DeepSeek reroutes every V4-Pro API request to V4.1-Flash from 14 September

    Build · September 22, 2026 · 1 publisher

  17. Luna lands at one tenth of Terra's price on both input and output tokens

    Build · September 22, 2026 · 1 publisher

  18. SpaceX prices Grok 4.7 at $4.69 a task on a benchmark it owns

    Product · September 21, 2026 · 1 publisher

  19. OpenRouter's P50 puts Mercury 2.5 at 440 tok/s against Inception's reported 1,107

    Build · September 21, 2026 · 1 publisher

  20. A model string one character off bills cached tokens at four times the rate

    Build · September 20, 2026 · 1 publisher

  21. Mystery model Union Alpha hit a billion tokens a minute before vanishing from listings and being revealed as Pareto

    Build · September 20, 2026 · 1 publisher

  22. Anthropic prices its newer Sonnet a third below Sonnet 4.5

    Build · September 20, 2026 · 1 publisher

  23. Meta charges 12.5 times more for Muse input tokens it promises not to train on

    Product · September 19, 2026 · 1 publisher

  24. A 90% cache-read discount takes 81% off a 10,000-token prompt's input line

    Build · September 17, 2026 · 1 publisher

  25. Every Qwen3.8-Omni-Flash workflow ends in text your own tools have to execute

    Build · September 17, 2026 · 1 publisher

  26. GLM 5.3's low effort setting misses five of 33 tasks its default gets right

    Build · September 15, 2026 · 1 publisher

  27. A 20-turn agent run bills 656,000 input tokens for 59,000 tokens of reading

    Build · September 15, 2026 · 1 publisher

  28. GPT-6 Astra lands in a different app depending on which ChatGPT plan you pay for

    Product · September 13, 2026 · 1 publisher

  29. PointFive's 230,000-token coding task produces a fivefold price gap between models

    Invest · September 12, 2026 · 1 publisher

  30. A top-level cache_control field moves the Claude cache breakpoint forward as the conversation grows

    Build · September 11, 2026 · 1 publisher

  31. FrugalGPT fits a fresh triage rule for every dataset and task it is tested on

    Build · September 10, 2026 · 1 publisher

  32. ARC Prize puts Astra 37 points below the score OpenAI led with

    Leadership · September 3, 2026 · 3 publishers

  33. DeepSeek's V4 Pro now bills seven hours a day at twice the off-peak rate

    Build · September 5, 2026 · 1 publisher

  34. Fable 5.1 doubles science benchmark score, cuts bug-hunt task time by 3.6 seconds

    Build · September 5, 2026 · 1 publisher

  35. Anthropic's 25% cheaper Fable 5.1 discounts one of six lines on the price sheet

    Product · September 3, 2026 · 1 publisher

  36. Anthropic bills Pro seats extra for the flagship model already in their picker

    Product · September 2, 2026 · 1 publisher

  37. Claude Opus 5 at $5/$25: the agent-loop math the rate card does not show

    Build · August 27, 2026 · 1 publisher

  38. GPT-5.6 ships as three models, and that makes model choice a deployment decision

    Build · August 18, 2026 · 1 publisher

  39. Four frontier models in four days, and the cheapest number in your agent plan has an expiry date

    Build · August 18, 2026 · 1 publisher

  40. DeepSeek's 12x cached-token rise ends the cheap-endpoint era for Chinese inference

    Invest · August 17, 2026 · 1 publisher