Anthropic released Sonnet 5.5 at $2 and $10 per million input and output tokens, half the Opus 5.5 rate. How much a buyer saves by moving work down a tier depends on tokens burned per task and on cache reads priced identically on both models.
Perspective Coverage
5 publishers
- Builder
- Builder 46%
- Operator
- Operator 38%
- Investor
- Investor 16%
Reality
- Evidence55
- Adoption35
- Hype gap+20
- Incentives70
- Confidence60
Microsoft released metadata from 301,026 GitHub Copilot agent sessions, covering 9.3 million LLM calls in one June week. Capacity planners now have real cache and token figures to test against, though the files measure resource use only and cannot show whether the output was any good.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap−5
- Incentives45
- Confidence60
One developer's month of Claude Code logs shows each break past the one-hour cache lifetime rewriting 125k-135k tokens at twice the base input price. What that costs a team depends on its cache timer and how big sessions grow before a break.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Cache-augmented generation costs about what retrieval does when the corpus is roughly 10 times the tokens retrieval would send, a dev.to analysis finds. Sparse traffic breaks the rule, because each query then pays the cache-write premium and caching becomes the most expensive option.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
OpenAI's GPT-6.1 Sol halves old Sol's cache-read rate to $0.10 per million tokens and requires the Responses API for tool calls. In one worked example the cut saves about 10%, and older agents only get that saving after their tool and reasoning fields are rewritten.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives35
- Confidence50
OpenAI's GPT-6.1 Sol halves cached-input pricing to $0.10 per million tokens and leaves standard rates at $2 and $10. Agents that resend long context collect the saving, while other buyers weigh gains shown mostly in OpenAI's own tests.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives65
- Confidence55
Opus 5.5 matched Opus 5 on two reasoning puzzles in The New Stack's tests at 43 to 69 percent lower cost. Both ran at default effort, medium on the new model and high on the old, so the saving a team sees depends on the effort level it pins.
Perspective Coverage
17 publishers
- Builder
- Builder 43%
- Operator
- Operator 33%
- Investor
- Investor 24%
Reality
- Evidence62
- Adoption48
- Hype gap+22
- Incentives58
- Confidence58
AWS Cost Anomaly Detection works from Cost Explorer data up to 24 hours old, so a dollar alarm on an agent fires after the money is spent. OpenTelemetry's GenAI spec has no cost attribute either, so teams must price each span from cache-split tokens and sum the trace tree.
Reality
- Evidence55
- Adoption20
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
OpenAI priced GPT-6 Sol at $2/$10 and Luna at $0.10/$0.50 per million tokens on September 22, half its GPT-5.6 promotional rates. Moving a job from Sol to Luna cuts its token rate by 95%, a bigger saving than the halving for any team whose work Luna can handle.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives65
- Confidence60
Jev's own gateway benchmark shows routing raised Opus 5 input tokens 61% on a Claude Code feature task, where the gateway can only hint at tools. Any saving depends on the task and on how many tools Claude Code sends the router each turn.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence40
Every off-peak rate sits above the old flat price, and Pro cache hits jumped roughly 6x. Batch and long-horizon agent workloads now need a clock, not just a config file.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+20
- Incentives40
- Confidence70
A 20% input and 33% output cut on GPT-5.6 Sol comes with a November 21 floor, while per-request regional routing makes data residency cheaper than plain global processing was in July.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+5
- Incentives50
- Confidence70
Opus 5.5 lists at twice Sonnet 5's price, yet in one developer's matched Claude Code runs it cost about 0.45 times as much per changed line. Agents re-send their whole context on every call, so fewer calls and fewer review rounds outweighed the higher token price.
Reality
- Evidence45
- Adoption8
- Hype gap+15
- Incentives
- Insufficient
- Confidence38
The Neuron gave it a browser, a Mac, a broken Blender install and an hour. What changed was how rarely it stopped to ask permission, which makes the next piece of work a harness problem rather than a prompt problem.
Perspective Coverage
14 publishers
- Builder
- Builder 42%
- Operator
- Operator 32%
- Investor
- Investor 26%
Reality
- Evidence50
- Adoption40
- Hype gap+25
- Incentives55
- Confidence55
A dev.to guide to GPT-5.6 pricing shows batch halving both token rates and a cheap-first cascade saving money until 9 in 10 calls escalate. Batch is opt-in and caching fails silently on short prefixes, so the default request often pays list price.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives75
- Confidence50
Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.
Perspective Coverage
5 publishers
- Builder
- Builder 38%
- Operator
- Operator 37%
- Investor
- Investor 25%
Reality
- Evidence55
- Adoption20
- Hype gap+25
- Incentives70
- Confidence60
Anthropic's new flagship lists at $4 and $20 per million tokens, a fifth under Opus 5, and cache reads drop 60 percent to $0.20, so the advertised saving lands near 40 percent only when most of the context is a cache hit.
Reality
- Evidence55
- Adoption45
- Hype gap+32
- Incentives75
- Confidence62
OpenAI's guide prices reused input tokens at a discount of up to 90 percent. It also says cached key-value states sit on individual machines, so an unchanged prefix can still miss the cache when routing overflows.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives75
- Confidence65
OpenAI's own benchmarks put GPT-6 Sol and Luna at half the list price and a fraction of a rival's cost per finished task. For defenders, the thing getting cheaper is autonomous tool calls into business systems.
Reality
- Evidence32
- Adoption38
- Hype gap+34
- Incentives80
- Confidence52
Every token price on Claude Opus 5.5 fell 20% except cache reads, which dropped 60% to 20 cents a million. For the advertised 40% saving to come from price alone, half a buyer's Opus 5 bill has to be cache reads.
Reality
- Evidence38
- Adoption24
- Hype gap+32
- Incentives78
- Confidence44
Earlier coverage
- Each plan-mode toggle under opusplan invalidates Claude Code's prompt cache
Build · September 23, 2026 · 1 publisher
- OpenAI cuts prices on new GPT-6 Sol and Luna models
Product · September 23, 2026 · 1 publisher
- DeepSeek reroutes every V4-Pro API request to V4.1-Flash from 14 September
Build · September 22, 2026 · 1 publisher
- Anthropic's worked example turns a 120,000-token conversation into 2.8 million billed input tokens
Invest · September 22, 2026 · 1 publisher
- A model string one character off bills cached tokens at four times the rate
Build · September 20, 2026 · 1 publisher
- Kimi K3 puts explicit prompt caching on Bedrock behind a 1,024-token minimum prefix
Build · September 18, 2026 · 1 publisher
- A 90% cache-read discount takes 81% off a 10,000-token prompt's input line
Build · September 17, 2026 · 1 publisher
- Ollama divides the whole prompt by the time it spent computing one token of it
Build · September 15, 2026 · 1 publisher
- Bedrock's cache write on the first request pulls the 90 percent discount down to 75
Build · September 15, 2026 · 1 publisher
- A 20-turn agent run bills 656,000 input tokens for 59,000 tokens of reading
Build · September 15, 2026 · 1 publisher
- DeepSeek prices cached agent input at $0.003 a million tokens off-peak
Product · September 12, 2026 · 1 publisher
- A timestamp in the system prompt turns prompt caching into a 25% surcharge
Build · September 12, 2026 · 1 publisher
- A top-level cache_control field moves the Claude cache breakpoint forward as the conversation grows
Build · September 11, 2026 · 1 publisher
- One API key per platform puts every customer's system prompt in the same cache namespace
Build · September 5, 2026 · 1 publisher
- Customer name placed ~1,000 tokens in coincided with a common 1,024-token cache floor in call logs
Build · September 4, 2026 · 1 publisher
- Anthropic's 25% cheaper Fable 5.1 discounts one of six lines on the price sheet
Product · September 3, 2026 · 1 publisher
- Compaction that cut tool output 38.4% pushed the bill up 6.8%
Build · September 2, 2026 · 1 publisher
- DeepSeek V4 moves the coding-model decision into the finance column
Build · September 1, 2026 · 1 publisher
- Per-PTU throughput spans 25x across three models in the same GPT-5.6 family
Build · August 30, 2026 · 1 publisher
- A semantic cache hit saves five times what a prompt cache hit saves
Build · August 29, 2026 · 1 publisher
- A cached prompt prefix repays its write premium on the second request
Build · August 28, 2026 · 1 publisher
- Anthropic's GA Files API re-bills the whole document on every request
Build · August 27, 2026 · 1 publisher
- Claude's limits are token meters on two clocks, and your open session is what drains them
Build · August 25, 2026 · 1 publisher
- Coding agents cost $4,125 a month because 73% of it is context you already sent
Build · August 23, 2026 · 1 publisher
- Tier the models; the validation boundary is the thing you are actually buying
Build · August 22, 2026 · 1 publisher
- Claude's prompt cache dies quietly in agent loops: the 20-block lookback nobody configures
Build · August 21, 2026 · 1 publisher
- Anthropic ships a cache differ, and concedes prompt caching was failing silently
Build · August 17, 2026 · 1 publisher
- DeepSeek's 12x cached-token rise ends the cheap-endpoint era for Chinese inference
Invest · August 17, 2026 · 1 publisher
- Prompt caching cuts agent API costs 41-80%, but only if tool results stay out of the cache
Build · August 16, 2026 · 1 publisher