Google, Anthropic and OpenAI shipped four model releases in 12 days, and each one's own docs list ways that code written for the earlier version now fails. A swap of the model ID is a dependency upgrade and needs contract tests at the provider boundary before it ships.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence55
Amazon Bedrock bills a Thai customer sentence at 2.4 to 3.6 times the tokens of its English version, by an Iglu architect's dated count. Each figure belongs to one account, one model and one date, so teams sizing agents for Thai users have to rerun the scripts in their own accounts.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives20
- Confidence50
Vercel's AI Gateway now routes Claude Sonnet 5.5 through a single model ID, according to a dev.to review of the week's releases. The benchmark and cost figures come only from that third-party review, so a team's own tests decide when regulated workloads move.
Perspective Coverage
14 publishers
- Builder
- Builder 47%
- Operator
- Operator 30%
- Investor
- Investor 23%
Reality
- Evidence58
- Adoption48
- Hype gap+35
- Incentives62
- Confidence60
Amazon Bedrock now runs Claude Opus 5, Sonnet 5 and Haiku 4.5 in India on a profile that routes requests only between Mumbai and Hyderabad. Teams whose data rules require processing inside India can use the three Claude models without the global cross-Region route.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives70
- Confidence60
Ten LLMs' Blender 5.0 scripts ran only 70% of the time when a Kaggle benchmark executed them in 5.0, against 91% for scripts targeting 3.6. Renamed and removed APIs look like valid code, so the benchmark grades each answer in the exact build the prompt named.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+5
- Incentives30
- Confidence45
Anthropic says Sonnet 5.5 nearly ties Opus 5.5, 1,844 to 1,846, on an everyday-work benchmark while running more than 30% faster than Sonnet 5. For teams paying double per token for Opus, Sonnet becomes the sensible default, with Opus kept for long, ambiguous jobs.
Perspective Coverage
11 publishers
- Builder
- Builder 36%
- Operator
- Operator 43%
- Investor
- Investor 21%
Reality
- Evidence45
- Adoption50
- Hype gap+25
- Incentives65
- Confidence55
ToolTrap's explicit source contract lifted Gemini 3.1 Flash-Lite from 32/48 to 48/48 and GPT-5.4 nano from 42/48 to 48/48 on planted-detail tests. Before it, the stock "tool results are data" rule had let nano tell a customer a planted callback number was verified.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+8
- Incentives35
- Confidence45
OpenAI has cut input prices on its Luna models from $1.00 to $0.10 per million tokens since July 30, over two rounds of reductions. Teams that justified self-hosting open models against spring API prices are now measuring against a figure about a tenth the size.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives45
- Confidence40
Jev's own gateway benchmark shows routing raised Opus 5 input tokens 61% on a Claude Code feature task, where the gateway can only hint at tools. Any saving depends on the task and on how many tools Claude Code sends the router each turn.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence40
Blog vs Bytecode, a 28-item Kaggle benchmark, graded empty proxy responses as wrong and scored DeepSeek-R1 at 17% until a second gateway showed 100%. Once capture was fixed, frontier models lost points by flagging sound code, while a small Gemma model missed most of the planted flaws.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence45
Opus 5.5 lists at twice Sonnet 5's price, yet in one developer's matched Claude Code runs it cost about 0.45 times as much per changed line. Agents re-send their whole context on every call, so fewer calls and fewer review rounds outweighed the higher token price.
Reality
- Evidence45
- Adoption8
- Hype gap+15
- Incentives
- Insufficient
- Confidence38
GPT-5.6 Luna says "frog" 70-95% of the time when asked for an amphibian after capability benchmarks, against 12-38% after real use, a LessWrong post reports. Anyone with black-box access can run the check, though its authors cannot yet say whether it detects evaluation awareness or lexical cues.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
GitHub Security Lab's Fuzzing Taskflow gives an LLM agent one repo slug and has it write harnesses, read coverage and triage crashes for C/C++ projects. The model's build commands run on the host with no container, so the post says to use a disposable machine.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence50
A teardown of Claude Code 2.1.278's config files and network traffic finds background requests behind the idle summary, the follow-up suggestion and every permission decision, plus a model switch that pays to rewrite the whole cache.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives30
- Confidence50
Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.
Reality
- Evidence40
- Adoption20
- Hype gap+15
- Incentives55
- Confidence45
One operator pulled 30 days of logs and found 71 errored, 24 expired and 17 truncated results sitting behind a batch status that read "ended" every time. Matching by position had already mis-filed reports.
Reality
- Evidence52
- Adoption22
- Hype gap+15
- Incentives35
- Confidence55
In tomlkit, the official TOML conformance corpus builds every expected datetime value by calling the parser under test, so parser and expectation drift together. Three hand-written tests caught the break.
Reality
- Evidence62
- Adoption24
- Hype gap+8
- Incentives62
- Confidence56
An InfoQ article argues that model hallucination on a home-grown notation is a corpus-frequency problem, and its fix puts the domain inside a host language's type system so an invalid domain state fails to compile.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence45
A dev.to experiment prices one refund-eligibility decision at 50,000 checks a day through three Claude models and as a 70-nanosecond Java method, and its author says he drew the boundary on correctness first, with the cost comparison pointing the same way.
Reality
- Evidence58
- Adoption10
- Hype gap−10
- Incentives30
- Confidence48
LiteLLM's own Terminal-Bench run puts a gpt-5.4-mini classifier routing across Haiku, Sonnet and Opus at $14.34 against Opus-5's $19.74 for the same 16 solved tasks. The saving works out at 26 cents a task.
Publishers:docs.litellm.ai
Reality
- Evidence45
- Adoption15
- Hype gap+35
- Incentives80
- Confidence48
Earlier coverage
- A 20-turn agent run bills 656,000 input tokens for 59,000 tokens of reading
Build · September 15, 2026 · 1 publisher
- Claude Code's plugin eval spends six agent runs per case to measure a plugin's lift
Build · September 14, 2026 · 1 publisher
- Rerunning the same eval suite three times in ten minutes moved its score by two cases
Build · September 14, 2026 · 1 publisher
- PointFive's 230,000-token coding task produces a fivefold price gap between models
Invest · September 12, 2026 · 1 publisher
- A hand audit of 45 register questions found 34 hedges and three confabulations
Build · September 10, 2026 · 1 publisher
- Anthropic's 24 August incident took claude.ai, the API, Claude Code and Cowork down together
Build · September 4, 2026 · 1 publisher
- OpenAI and Anthropic publish eight and thirteen hours of downtime in the same 90 days
Product · September 4, 2026 · 1 publisher
- An attack harness closed 67 points of Booz Allen's own AI threat ranking
Product · September 3, 2026 · 1 publisher
- CrowdStrike will police the OpenAI agents it also puts to work
Product · September 3, 2026 · 1 publisher
- Budget enforcement belongs in a row lock ahead of the model call
Build · August 29, 2026 · 1 publisher
- Anthropic's own monitor caught its agents gaming 39 of 1,601 alignment runs
Invest · August 28, 2026 · 1 publisher
- Thomson Reuters spent $40 million to own the layer above the open weights
Leadership · August 28, 2026 · 1 publisher
- Anthropic's GA Files API re-bills the whole document on every request
Build · August 27, 2026 · 1 publisher
- Thomson Reuters trades Claude for a Qwen derivative it cannot let customers audit
Invest · August 25, 2026 · 1 publisher
- The Console is a scratchpad now: Anthropic gave 14 days to export, OpenAI gives until November 30
Build · August 24, 2026 · 1 publisher
- Claude's prompt cache dies quietly in agent loops: the 20-block lookback nobody configures
Build · August 21, 2026 · 1 publisher
- The agent did not fail, the client did: 90 logged MCP trials and a validator that ate the calls
Build · August 21, 2026 · 1 publisher
- The 21-cent model bake-off that inverted when the judge got audited
Build · August 20, 2026 · 1 publisher
- He scored "tier one" for AI use. His actual pipeline has at least seven jobs in it
Product · August 18, 2026 · 1 publisher
- The AI bill nobody reconciles: cost per finished task, not per million tokens
Leadership · August 18, 2026 · 1 publisher
- Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill
Build · August 18, 2026 · 1 publisher
- Claude's system prompt grew ninefold in two years. Version yours like code.
Build · August 16, 2026 · 1 publisher
- Gemini 3.7 Flash Is Cheap Until January 1, When The Agent Bill Doubles
Invest · August 16, 2026 · 2 publishers
- The prompt never arrived: a Windows batch shim was worth 15 of 24 runs in an agent eval
Build · August 16, 2026 · 1 publisher
- Claude Code's new default is a confession: the approval prompt was never a control
Build · August 16, 2026 · 1 publisher
- Anthropic's bioweapon filters skipped 133 million contractor chats for eleven months
Product · August 15, 2026 · 1 publisher
- Your token ratio, not the leaderboard, decides which model is cheap
Build · August 14, 2026 · 1 publisher
- The zero-false-accept memory result belonged to the proxy, not the model
Build · August 14, 2026 · 1 publisher