Anthropic released Sonnet 5.5 at $2 and $10 per million input and output tokens, half the Opus 5.5 rate. How much a buyer saves by moving work down a tier depends on tokens burned per task and on cache reads priced identically on both models.
Perspective Coverage
5 publishers
- Builder
- Builder 46%
- Operator
- Operator 38%
- Investor
- Investor 16%
Reality
- Evidence55
- Adoption35
- Hype gap+20
- Incentives70
- Confidence60
NVIDIA released SoL-Pi, an MIT-licensed Pi extension with four token-saving harness mechanisms that an optimizer agent found. A dev.to review puts the saving near one third on long sessions, bought with a few lost solves.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence45
Vercel's AI Gateway now routes Claude Sonnet 5.5 through a single model ID, according to a dev.to review of the week's releases. The benchmark and cost figures come only from that third-party review, so a team's own tests decide when regulated workloads move.
Perspective Coverage
14 publishers
- Builder
- Builder 47%
- Operator
- Operator 30%
- Investor
- Investor 23%
Reality
- Evidence58
- Adoption48
- Hype gap+35
- Incentives62
- Confidence60
Google Cloud AI Research has released RRSI, an Apache 2.0 tool whose self-rewriting agent harness lifted Terminal-Bench 2.1 scores from 74.2% to 80.2%. Any team can use it commercially, though on tasks the agent never trained against the reported gain falls to 4.7 points.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence40
Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 30%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption30
- Hype gap+25
- Incentives70
- Confidence58
Claude Opus 5.5 matched Fable 5.1 on every hidden test in two New Stack coding trials, at $0.75 and $1.42 per run against $1.50 and $1.96. It needed more tokens and more minutes to get there, so the saving holds only for tasks that resemble these.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
Amazon Bedrock now serves xAI's 500K-token-context Grok 4.7 through the Responses, Chat Completions and Converse APIs. Trying it from an existing client takes little code, though Artificial Analysis found its gains cost about twice the output tokens per task.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives70
- Confidence60
Anthropic says Sonnet 5.5 nearly ties Opus 5.5, 1,844 to 1,846, on an everyday-work benchmark while running more than 30% faster than Sonnet 5. For teams paying double per token for Opus, Sonnet becomes the sensible default, with Opus kept for long, ambiguous jobs.
Perspective Coverage
11 publishers
- Builder
- Builder 36%
- Operator
- Operator 43%
- Investor
- Investor 21%
Reality
- Evidence45
- Adoption50
- Hype gap+25
- Incentives65
- Confidence55
Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
Perspective Coverage
4 publishers
- Builder
- Builder 51%
- Operator
- Operator 31%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence58
Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.
Perspective Coverage
5 publishers
- Builder
- Builder 58%
- Operator
- Operator 33%
- Investor
- Investor 9%
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives70
- Confidence55
Nvidia's SoL-Pi, an automated search over coding-agent harnesses, cut token use 44.7 to 49 percent at scores close to the Pi baseline. The gains were measured on 40 held-out tasks with a search fitted to one model, so they carry over only as far as a team's workload resembles that setup.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence50
The top four slots on Artificial Analysis are the advertisement. The line item Anthropic actually moved is the one that scales with how long an agent runs, and its own savings estimate backs out that share at about 60 percent.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives62
- Confidence60
The replay only works on runs Warp's infrastructure already recorded, and correctness is graded by a judge model against a rubric you write. Warp's own 30-task bake-off cost $2,130.57. It finished in under four hours.
Publishers:runtimewire.com · warp.dev Reality
- Evidence40
- Adoption15
- Hype gap+35
- Incentives80
- Confidence55
Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.
Perspective Coverage
8 publishers
- Builder
- Builder 53%
- Operator
- Operator 29%
- Investor
- Investor 18%
Reality
- Evidence58
- Adoption35
- Hype gap+22
- Incentives72
- Confidence62
Meta's fourth Muse Spark in five months gains four points on Artificial Analysis' Intelligence Index, almost entirely in agentic tests, while two scores fall and the tier Meta benchmarked against rivals is still a limited preview.
Perspective Coverage
6 publishers
- Builder
- Builder 52%
- Operator
- Operator 26%
- Investor
- Investor 22%
Reality
- Evidence68
- Adoption25
- Hype gap+30
- Incentives65
- Confidence70
OpenAI's spec lets gpt-6-astra take 922,000 input tokens, but requests above 272,000 move to higher long-context rates. Provisioning a model that drives a desktop, a shell and MCP servers starts with that threshold.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence40
Grok 4.7 arrived on Monday after five walked-back timelines, with 40% more parameters and Grok 4.6's list price intact. Cursor's own cost chart still puts its price per task above GPT-6 Astra and Claude Sonnet 5.
Perspective Coverage
3 publishers
- Builder
- Builder 42%
- Operator
- Operator 27%
- Investor
- Investor 31%
Reality
- Evidence45
- Adoption35
- Hype gap+15
- Incentives70
- Confidence55
Anthropic cut Opus 5.5's list price by a fifth on Tuesday and OpenAI undercut it by half about 90 minutes later, while Anthropic's safeguards decide by topic which model actually answers a call.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+25
- Incentives65
- Confidence60
Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.
Perspective Coverage
5 publishers
- Builder
- Builder 38%
- Operator
- Operator 37%
- Investor
- Investor 25%
Reality
- Evidence55
- Adoption20
- Hype gap+25
- Incentives70
- Confidence60
Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.
Perspective Coverage
7 publishers
- Builder
- Builder 41%
- Operator
- Operator 38%
- Investor
- Investor 21%
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives60
- Confidence60
Earlier coverage
- Anthropic cuts Opus 5.5 prices 20% on tokens, 60% on cache reads, citing fewer tokens burned for 40% total savings
Invest · September 23, 2026 · 1 publisher
- Grok 4.7 buys 5.9 points of CursorBench at Grok 4.6's token price
Invest · September 22, 2026 · 1 publisher
- Fixing the deployment target splits the flash-tier coding leaderboard into three winners
Build · September 21, 2026 · 1 publisher
- Strands Harness keeps five subsystems local and routes one call to Bedrock
Build · September 21, 2026 · 1 publisher
- A Berkeley scanning agent scores 100% on five AI agent benchmarks without solving a task
Science · September 20, 2026 · 1 publisher
- Anthropic prices its newer Sonnet a third below Sonnet 4.5
Build · September 20, 2026 · 1 publisher
- Swapping the harness under a fixed model kept pass rate within 8 points on 50 Terminal-Bench Pro tasks
Build · September 19, 2026 · 1 publisher
- Artificial Analysis retries a provider safety error ten times before scoring the attempt zero
Build · September 19, 2026 · 1 publisher
- Rule-based elision before summarization led on efficiency among context-management strategies
Leadership · September 19, 2026 · 1 publisher
- Databricks finds the harness swings coding-agent cost more than the model does
Product · September 18, 2026 · 1 publisher
- Real-SWE licenses private production codebases to score coding agents on real business tasks
Build · September 18, 2026 · 1 publisher
- Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test
Build · September 17, 2026 · 1 publisher
- Fireworks' own DeepSWE numbers put four coding models inside the noise band
Product · September 17, 2026 · 1 publisher
- LangChain's Deep Agents offloads oversized tool output to a filesystem at 20,000 tokens
Build · September 16, 2026 · 1 publisher
- Arize's cheapest model per finished task reliably solves only a fifth of the benchmark
Leadership · September 15, 2026 · 1 publisher
- Amodei's pacing plan would put outside evaluators inside Anthropic with the right to publish
Leadership · September 15, 2026 · 1 publisher
- DeepSeek's smallest model beats its own 28-day-old flagship on seven of eight shared scores
Invest · September 14, 2026 · 1 publisher
- DeepSeek's V4.1-Flash reads a million-token prompt on 8B active parameters
Leadership · September 12, 2026 · 1 publisher
- GLM-5.3-Flash buys seven retries for the price of one Kimi K3 call
Build · September 11, 2026 · 1 publisher
- DeepSeek reroutes V4-Pro API traffic to a smaller model on September 14
Product · September 11, 2026 · 1 publisher
- ByteDance's self-evolved agent harnesses gain 3.11 held-out points inside a 4.75-point noise band
Build · September 10, 2026 · 1 publisher
- Cognition put a cost penalty inside SWE-2's reinforcement-learning objective
Build · September 10, 2026 · 1 publisher
- DeepSeek's V4 preview cuts million-token KV cache to a tenth of V3.2's
Leadership · September 8, 2026 · 1 publisher
- Huang tucks 400,000 incoming GPUs into three-word "AGI has arrived" post
Product · September 7, 2026 · 1 publisher
- Abliteration.ai rents a refusal-stripped GLM-5.3 for five dollars a million tokens
Build · September 6, 2026 · 1 publisher
- Post-training alone took GLM-5.3 from 4.6 to 28.3 on Terminal-Bench 3.0
Build · August 28, 2026 · 8 publishers
- GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor
Leadership · September 5, 2026 · 1 publisher
- Four leaderboards, four denominators: what you buy when you standardize on a coding agent
Product · August 25, 2026 · 1 publisher
- Harness-Bench makes the scaffold a measured variable, and model-only scores lose meaning
Build · August 23, 2026 · 1 publisher
- GLM-5.3 changed nothing but the training environments. That is the whole test.
Build · August 19, 2026 · 3 publishers
- Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist
Leadership · August 18, 2026 · 1 publisher
- Three frontier launches in a day, all pitched on price. Open weights set the ceiling.
Build · August 14, 2026 · 4 publishers