AWS's project spending limits pause a project at its cap and permanently delete its data after 90 days paused without action. The cap bounds what a runaway agent experiment can bill, so someone has to own recovery and keep independent backups.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence55
Agent runtimes should refuse unit 501 of a 500-unit budget before the call runs, a dev.to post argues. The new AWS and Google Cloud spend caps then become the backstop behind a tighter limit in the agent's own call path.
Reality
- Evidence55
- Adoption20
- Hype gap+15
- Incentives
- Insufficient
- Confidence60
AWS added a per-project spend limit on September 16, 2026 that pauses service at the monthly cap, following Google Cloud's July launch of Spend Caps. A cap that checks each request stops at once, while one built on lagging billing data keeps charging until a function fires.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence30
Diogo Almeida's TypeSafe put Jev into early access on 15 September with $40m from DCVC, and it sells a calibrated confidence number on every answer as the thing that makes automation possible, with the evaluations behind that claim built in-house.
Reality
- Evidence45
- Adoption45
- Hype gap+25
- Incentives60
- Confidence55
Prism's ternary Bonsai 2 27B ran at two to three tokens a second on a CPU-only Hetzner VPS in a dev.to test, against Simon Willison's 20 to 44 on a Mac. The sub-6GB file fits a 16GB box easily, but at that speed it only suits batch jobs.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence45
Shopify is moving every one of its apps off React Native to native Swift and Kotlin, and rebuilt its Shop app that way in 12 weeks. The switch leaves three Shopify-maintained React Native libraries, including the widely used FlashList, archived or without a maintainer.
Reality
- Evidence48
- Adoption63
- Hype gap+14
- Incentives55
- Confidence42
Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
Vercel Labs' experimental ScriptC compiles TypeScript to native binaries that started in 1.78ms against Node's 61.78ms in one benchmark. The gain holds for short, statically typed programs, while framework code and sustained compute both ran slower than Bun or Node.
Reality
- Evidence55
- Adoption15
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence50
TypeSafe's Jev answers typed questions with floats and probabilities. An invoice pipeline that handed it classification and catalogue selection still needs a generative model for field extraction and for the note a human reads.
Perspective Coverage
14 publishers
- Builder
- Builder 53%
- Operator
- Operator 31%
- Investor
- Investor 16%
Reality
- Evidence55
- Adoption40
- Hype gap+30
- Incentives65
- Confidence55
Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
Perspective Coverage
4 publishers
- Builder
- Builder 51%
- Operator
- Operator 31%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence58
Enterprises buy the cheapest model that clears their bar. On Ramp's July billing data, that leaves Anthropic's flagship with about an eighth of its maker's platform spend.
Reality
- Evidence55
- Adoption25
- Hype gap+30
- Incentives40
- Confidence55
Hy4 preview carries 2.6 times Hy3's parameters and 3.9 times its context window. Tencent is giving the weights away. That means the build column in next quarter's model budget gets priced in accelerator memory rather than in tokens.
Publishers:simonwillison.net · tencent.com Reality
- Evidence62
- Adoption18
- Hype gap+30
- Incentives65
- Confidence60
Simon Willison's teardown finds a cloud container whose outbound domain list appears open by default, plus a desktop build that runs programs on the employee's machine. Those are two provisioning reviews with different threat models.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
OpenAI's write-up gives the token counts, the agent count and the verification time for its Navier-Stokes result. The verification time is the number that decides whether the method transfers to anyone else's workload.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence55
Six engineers shipped a fully native Shop app in 12 weeks, but the first attempt at handing the React Native code to a model produced what Shopify's engineering director called slop. Helix chunks each screen instead.
Perspective Coverage
5 publishers
- Builder
- Builder 56%
- Operator
- Operator 33%
- Investor
- Investor 11%
Reality
- Evidence70
- Adoption55
- Hype gap+20
- Incentives50
- Confidence68
Three of the four authors of last week's wiki-agent report say an OpenAI swarm very likely published the hundreds of packages that hit RubyGems on 12 May, and their strongest evidence is a retrieval trick the wiki agents also used.
Perspective Coverage
8 publishers
- Builder
- Builder 36%
- Operator
- Operator 39%
- Investor
- Investor 25%
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence65
OpenAI disclosed the incident on September 16 under a framework it polices itself. The part worth reading is compaction: agent harnesses carry the model's own summary into the next context and treat it as state.
Perspective Coverage
13 publishers
- Builder
- Builder 43%
- Operator
- Operator 38%
- Investor
- Investor 19%
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence58
Anthropic and OpenAI shipped cheaper model tiers minutes apart on Tuesday. OpenAI halved Sol's posted token prices, and the only per-task cost comparison between the two labs so far comes from OpenAI itself.
Perspective Coverage
16 publishers
- Builder
- Builder 29%
- Operator
- Operator 35%
- Investor
- Investor 36%
Reality
- Evidence62
- Adoption25
- Hype gap+25
- Incentives70
- Confidence60
AWS closed Unit 42's AgentCore finding as informative and put tool scoping on the customer. Every step the agent took to leak the token was a capability someone granted it on purpose. That moves the control into the grant list.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+18
- Incentives55
- Confidence45
Earlier coverage
- A spreadsheet on Hugging Face tested whether its processor could reach Azure metadata
Product · September 18, 2026 · 1 publisher
- Apollo's Watcher escalates a flagged agent action to a bigger AI before any human sees it
Product · September 17, 2026 · 1 publisher
- Security practitioners put logging and permissions ahead of Amodei's audit plan
Product · September 16, 2026 · 1 publisher
- Alibaba ships the Qwen4 architecture as open weights before the flagship exists
Build · August 28, 2026 · 5 publishers
- Claude's Gmail agent turns one approval toggle into your whole outbound policy
Product · September 6, 2026 · 1 publisher
- Fable 5.1's 52.6% science score arrives on a benchmark that was five days old
Build · September 1, 2026 · 1 publisher
- Ramp's July card data puts Opus 4.8 at 3.5 times Claude Fable's spend share
Product · August 30, 2026 · 1 publisher
- A model documenting a retry wrapper hands you tenacity's parameters
Build · August 27, 2026 · 1 publisher
- Fable 5 at $50 per million output tokens turns model routing into a budget line
Build · August 23, 2026 · 2 publishers
- Mojo's compiler went Apache 2.0 fifty-five days after Qualcomm's $3.92bn deal
Build · August 21, 2026 · 1 publisher
- Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens
Build · August 21, 2026 · 1 publisher
- The sandbox is the product: what user-generated features actually require
Build · August 19, 2026 · 2 publishers
- A 27B laptop model scores like a rented one, and thinks three times as hard to do it
Product · August 19, 2026 · 1 publisher
- Claude's system prompt grew ninefold in two years. Version yours like code.
Build · August 16, 2026 · 1 publisher
- A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo
Build · August 15, 2026 · 1 publisher