Cerebras says splitting inference stages across chip types gave 5x more throughput from the same number of its systems without slowing token generation. Because the count covers only Cerebras hardware, the figure does not yet show what a mixed-chip fleet costs per unit of work.
Reality
- Evidence30
- Adoption20
- Hype gap+30
- Incentives75
- Confidence40
Cognition says its SWE-2 model produced up to 4.8 times more total token throughput on Nvidia's Vera Rubin NVL72 than on GB200, in a test on CoreWeave. The company-run result points to more agent capacity per system, but whether long coding tasks get cheaper depends on what the new systems cost per hour.
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives85
- Confidence60
Armin Ronacher let GPT-6 Astra code alone for 35 hours, spending about $1,200 on a net 75,000 lines of "absolutely nothing of value". His traces show habits from training that pays for finished tasks, and they reach the code when nobody reviews it.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives35
- Confidence50
DeepSeek 4.1 Flash finished a metered Ship-Bench build for $15.04 in per-token fees and scored lowest at code review. Developers who run more than about one and a third full builds a month would still pay less on a $20 seat.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.
Perspective Coverage
5 publishers
- Builder
- Builder 58%
- Operator
- Operator 33%
- Investor
- Investor 9%
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives70
- Confidence55
Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.
Perspective Coverage
8 publishers
- Builder
- Builder 53%
- Operator
- Operator 29%
- Investor
- Investor 18%
Reality
- Evidence58
- Adoption35
- Hype gap+22
- Incentives72
- Confidence62
Meta's fourth Muse Spark in five months gains four points on Artificial Analysis' Intelligence Index, almost entirely in agentic tests, while two scores fall and the tier Meta benchmarked against rivals is still a limited preview.
Perspective Coverage
6 publishers
- Builder
- Builder 52%
- Operator
- Operator 26%
- Investor
- Investor 22%
Reality
- Evidence68
- Adoption25
- Hype gap+30
- Incentives65
- Confidence70
Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.
Perspective Coverage
5 publishers
- Builder
- Builder 38%
- Operator
- Operator 37%
- Investor
- Investor 25%
Reality
- Evidence55
- Adoption20
- Hype gap+25
- Incentives70
- Confidence60
Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.
Perspective Coverage
7 publishers
- Builder
- Builder 41%
- Operator
- Operator 38%
- Investor
- Investor 21%
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives60
- Confidence60
Seat prices do not change, but the allowance sitting on top of each seat becomes a dollar balance of tokens, and the fallback to a cheaper model when that allowance runs out goes away on the same day.
Reality
- Evidence72
- Adoption55
- Hype gap+20
- Incentives75
- Confidence70
Fabian Hedin, Lovable's CTO, said two-thirds of the Fortune 500 use the tool because staff found it themselves, and put current revenue at $600m. The permissions work starts after the apps are running.
Reality
- Evidence32
- Adoption52
- Hype gap+36
- Incentives82
- Confidence46
The multipliers in circulation for how far AI seats are underpriced run from five times to twenty, and all of them rest on an unpublished per-seat token count. The API rate card is the part you can check.
Reality
- Evidence40
- Adoption35
- Hype gap+35
- Incentives58
- Confidence38
Anthropic's new flagship lists at $4 and $20 per million tokens, a fifth under Opus 5, and cache reads drop 60 percent to $0.20, so the advertised saving lands near 40 percent only when most of the context is a cache hit.
Reality
- Evidence55
- Adoption45
- Hype gap+32
- Incentives75
- Confidence62
A study of eight frontier models on SWE-bench Verified puts agentic coding at roughly 1,000 times the token cost of code chat, dominated by input, with the models' own pre-run estimates correlating no better than 0.39.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence55
Every token price on Claude Opus 5.5 fell 20% except cache reads, which dropped 60% to 20 cents a million. For the advertised 40% saving to come from price alone, half a buyer's Opus 5 bill has to be cache reads.
Reality
- Evidence38
- Adoption24
- Hype gap+32
- Incentives78
- Confidence44
Sol lists at $2 and $10 per million tokens and Luna at $0.10 and $0.50. The per-task savings OpenAI published come mostly from the lower price, and the cheaper tier scores below its predecessor on computer use.
Publishers:forkast.news · openai.com Reality
- Evidence55
- Adoption32
- Hype gap+30
- Incentives82
- Confidence62
Anthropic's cost breakdown for Opus 5.5 walks through a 40-turn coding task in which nine tenths of the input is resent conversation billed as cache reads. Cutting that line 60% takes the task's input bill down 39% on a 20% per-token cut.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives78
- Confidence62
Worker Previews give a branch its own URL, configuration and state, and isolating the Durable Object namespace is what keeps a preview's writes away from the single instance that owns production's storage.
Reality
- Evidence42
- Adoption18
- Hype gap+20
- Incentives85
- Confidence56
Across twelve ORBITAL tasks, eight failed a browser check at least once, and the verifier's own triage marked six of them automation bugs. Those six were cleared by override, two of them with per-case approval.
Reality
- Evidence45
- Adoption12
- Hype gap−10
- Incentives52
- Confidence45
GLM-5.3-Flash leads agentic terminal work, DeepSeek V4 Flash is billed as the cheapest per token, and a 2.52B MiniCPM5-2B runs locally under Apache 2.0. The comparison flags most of those numbers as vendor-reported.
Reality
- Evidence34
- Adoption27
- Hype gap+26
- Incentives58
- Confidence41
Earlier coverage
- Anthropic prices its newer Sonnet a third below Sonnet 4.5
Build · September 20, 2026 · 1 publisher
- A compliance agent reads its permission table and stops one transition short of signing
Build · September 19, 2026 · 1 publisher
- Thompson credits the harness for the agent jump that retired his bubble call
Build · September 19, 2026 · 1 publisher
- Lauren Tan's 2,000 pull requests a month rest on an app that fits in one process
Build · September 19, 2026 · 1 publisher
- Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test
Build · September 17, 2026 · 1 publisher
- OpenAI's improvement loop compiles five traced runs into a rerunnable Promptfoo gate
Build · September 15, 2026 · 1 publisher
- A four-hour Claude Code session landed 42 unreviewed commits on main
Build · September 14, 2026 · 1 publisher
- Anthropic takes one-cent warrants on 51 million shares of Truth Social's host
Invest · September 14, 2026 · 1 publisher
- A zero border radius rule caused more arguments with Claude Code than the 1C sync did
Build · September 13, 2026 · 1 publisher
- Parlotype's localization gate landed before roughly 450 keys left the markup
Build · September 11, 2026 · 1 publisher
- Overnight laptop runs took over most of one Rust developer's Opus coding work
Build · September 10, 2026 · 1 publisher
- Anthropic broke an agent ceiling by making "is this design good?" a gradable question
Leadership · September 10, 2026 · 1 publisher
- Anthropic sold about 8 percent of itself for $30 billion
Leadership · September 6, 2026 · 1 publisher
- Anthropic writes the agent handoff into the repository instead of the context window
Leadership · September 6, 2026 · 1 publisher
- Fable 5.1 doubles science benchmark score, cuts bug-hunt task time by 3.6 seconds
Build · September 5, 2026 · 1 publisher
- A $28 agent run swapped BCD for base-2^64 limbs and built its own oracle
Build · September 5, 2026 · 1 publisher
- Anthropic's own playbook moves the software bottleneck into the review queue
Product · September 3, 2026 · 1 publisher
- One Find and Replace task emptied a five-hour Codex window in 24 minutes
Build · August 29, 2026 · 1 publisher
- A 107-page specification turned Claude Code into three native Task Managers
Build · August 28, 2026 · 1 publisher
- Anthropic ships a price dial with its new model, and that is now the buying decision
Leadership · August 26, 2026 · 1 publisher
- JetBrains asked 15,000 developers how much code agents write. The answers add up to 112 percent
Build · August 26, 2026 · 1 publisher
- The $200 AI seat is a subsidy: ceiling users pull 40x to 70x what they pay
Invest · August 23, 2026 · 1 publisher
- A CEO's $1,000 weekend, and the auto-renew setting that made it possible
Invest · August 22, 2026 · 1 publisher
- Slack Code makes coding agents taggable teammates. Your merge-approval policy is now overdue
Product · August 20, 2026 · 4 publishers
- GLM-5.3 changed nothing but the training environments. That is the whole test.
Build · August 19, 2026 · 3 publishers