Google Cloud AI Research's RRSI lifted agent scores up to 4.7 points on five unseen benchmarks by capping how far a harness can rewrite itself. Its guardrails cost points on the tuning tasks, a trade worth making for teams that need harness gains to hold on new work.
Reality
- Evidence45
- Adoption8
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence40
Andon Labs opened Pion, its platform for agent-run companies, as a research preview on September 14, with its own agent-run store and cafe still losing money. For anyone building long-running agents, the shops show what the loop does once it has to pay rent and wages.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence40
A dev.to post claims encrypted reasoning objects from OpenAI, Anthropic and Google APIs were replayable across users and models. The disclosure is thin, but the storage habit it exposes is yours.
Reality
- Evidence66
- Adoption45
- Hype gap+20
- Incentives55
- Confidence60
Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.
Perspective Coverage
5 publishers
- Builder
- Builder 58%
- Operator
- Operator 33%
- Investor
- Investor 9%
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives70
- Confidence55
Claude Code got an HP Laser 1008a printing from an Apple Silicon Mac in about four hours. The MIT-licensed result works, and it leaves someone holding a root daemon and an always-on Linux VM.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence60
Claude agents designed 1,440 protein binders and 354 bound. The interesting part is that the prompts, provenance and every measurement went out with the number.
Reality
- Evidence60
- Adoption30
- Hype gap+12
- Incentives68
- Confidence62
Claude Opus 4.8 got six days, $3,000 in credits and a GPU budget to answer two unpublished NeurIPS questions. The papers' original authors graded the output and rejected both.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
V4-Flash-Vision-Exp is live on DeepSeek's API and costs a fraction of Anthropic's price. The vendor's own table shows it trailing on eight of eleven tests, including a 12-point gap on repository work.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence55
A flaw in NASA's open-source AIT-GUI lets unauthenticated requests reach spacecraft command routes, and Cycode says a malicious web page can deliver them through an operator's browser.
Perspective Coverage
4 publishers
- Builder
- Builder 43%
- Operator
- Operator 50%
- Investor
- Investor 7%
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap+30
- Incentives45
- Confidence65
Enterprises buy the cheapest model that clears their bar. On Ramp's July billing data, that leaves Anthropic's flagship with about an eighth of its maker's platform spend.
Reality
- Evidence55
- Adoption25
- Hype gap+30
- Incentives40
- Confidence55
Anthropic cut Opus 5.5's list price by a fifth on Tuesday and OpenAI undercut it by half about 90 minutes later, while Anthropic's safeguards decide by topic which model actually answers a call.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+25
- Incentives65
- Confidence60
Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.
Perspective Coverage
7 publishers
- Builder
- Builder 41%
- Operator
- Operator 38%
- Investor
- Investor 21%
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives60
- Confidence60
Five frontier LLMs with web search disagreed on 63% of 997 claims users sent to Lenz.io for checking. They also rated themselves 9 or 10 out of 10 on 76% of answers, so one model's confidence says little about whether the others would agree.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence55
Anthropic's help center says the safety classifiers on Opus 5 and 5.5 check memory, connector output, web results and files as well as the prompt. A fallback can therefore come from any of those.
Reality
- Evidence58
- Adoption42
- Hype gap−18
- Incentives62
- Confidence62
Automated harness evolution keeps whatever edits raise its own benchmark score, so the harness ends up fitted to the eval set. Google Cloud AI Research answers with five regularizers lifted from supervised learning.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+35
- Incentives45
- Confidence30
Tuesday's model releases from Anthropic and OpenAI came with audit numbers that move in both directions at once. For anyone granting an agent tokens or tooling, a version bump means re-running the injection tests.
Reality
- Evidence54
- Adoption30
- Hype gap+18
- Incentives79
- Confidence44
Commerce told Anthropic on June 12 that any foreign national, anywhere, needs a BIS license to use Fable 5 or Mythos 5. Anthropic disabled both models for every customer to comply, and the letter has not been made public.
Publishers:anthropic.com · csis.org · natlawreview.com Perspective Coverage
3 publishers
- Builder
- Builder 32%
- Operator
- Operator 41%
- Investor
- Investor 27%
Reality
- Evidence61
- Adoption77
- Hype gap+9
- Incentives74
- Confidence66
Every token price on Claude Opus 5.5 fell 20% except cache reads, which dropped 60% to 20 cents a million. For the advertised 40% saving to come from price alone, half a buyer's Opus 5 bill has to be cache reads.
Reality
- Evidence38
- Adoption24
- Hype gap+32
- Incentives78
- Confidence44
Of the ten code migrations Anthropic says finished in a month, only one carries a dollar figure: 5.9 billion input tokens and 690 million output tokens, a million lines of Rust, about 16.5 cents a line.
Reality
- Evidence45
- Adoption55
- Hype gap+30
- Incentives88
- Confidence50
Earlier coverage
- Filling GLM-5.3-Flash's million-token window costs three times its per-task benchmark price
Build · September 20, 2026 · 2 publishers
- Export controls took Anthropic's newest models offline three days after launch
Leadership · September 16, 2026 · 1 publisher
- LangChain drops about 4,000 base input tokens from every default Deep Agents turn
Build · September 15, 2026 · 1 publisher
- A three-agent CrewAI run spent 44.6 of its 118 seconds inside coworker tool calls
Build · September 14, 2026 · 1 publisher
- DeepSeek's smallest model beats its own 28-day-old flagship on seven of eight shared scores
Invest · September 14, 2026 · 1 publisher
- Andon Labs' Pion hands one agent the email, phone, banking and cards of a whole business
Build · September 14, 2026 · 1 publisher
- Anthropic's most capable model built working exploits from 16 of 39 published patches
Leadership · September 11, 2026 · 1 publisher
- Stripped GLM-5.3-Flash weights show what Z.ai's MIT license permits
Build · September 8, 2026 · 1 publisher
- GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor
Leadership · September 5, 2026 · 1 publisher
- Three lines of context in the LSP reply cut follow-up file reads from 15.2 to 3.2
Build · September 3, 2026 · 1 publisher
- An attack harness closed 67 points of Booz Allen's own AI threat ranking
Product · September 3, 2026 · 1 publisher
- Pandex hooked a Fortune 500 agent four minutes after claiming a package name from llms.txt
Build · September 2, 2026 · 1 publisher
- Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8
Build · August 31, 2026 · 1 publisher
- Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test
Build · August 31, 2026 · 2 publishers
- Microsoft retired four models from the Foundry router under every deployment left on defaults
Build · August 31, 2026 · 1 publisher
- Ramp's July card data puts Opus 4.8 at 3.5 times Claude Fable's spend share
Product · August 30, 2026 · 1 publisher
- Open weights take 29% of gateway tokens on a twenty-fifth of the dollars
Invest · August 30, 2026 · 1 publisher
- Anthropic's own monitor caught its agents gaming 39 of 1,601 alignment runs
Invest · August 28, 2026 · 1 publisher
- Thomson Reuters spent $40 million to own the layer above the open weights
Leadership · August 28, 2026 · 1 publisher
- Same price, cheaper fast mode: Opus 4.8 argues on unit economics
Leadership · August 26, 2026 · 1 publisher
- Ox Alpha was GLM-5.3-Flash, and the number that decides displacement is 18 billion
Product · August 26, 2026 · 1 publisher
- Anthropic ships a price dial with its new model, and that is now the buying decision
Leadership · August 26, 2026 · 1 publisher
- Google's legal AI bundle lands a day after a $40M model, and the connector list tells you why
Build · August 26, 2026 · 2 publishers
- Grayscale's ZCSH puts Zcash in brokerage accounts, and none of its ZEC is shielded
Invest · August 25, 2026 · 3 publishers
- Claude Code tells you the model, not the culprit: 106 lines of shell to name the Skill
Build · August 25, 2026 · 1 publisher
- Thomson Reuters trades Claude for a Qwen derivative it cannot let customers audit
Invest · August 25, 2026 · 1 publisher
- App factory or agent fleet manager: the fork is whose rate limit stops the work
Build · August 24, 2026 · 1 publisher
- Thomson Reuters priced the middle path at $40M, and still pays Anthropic
Build · August 24, 2026 · 4 publishers
- Four Claude models, four surfaces, one incident: tier fallback is inside the blast radius
Product · August 24, 2026 · 1 publisher
- The AI boss forgot its own handbook, and humans had to hand it back
Build · August 23, 2026 · 1 publisher
- Tier the models; the validation boundary is the thing you are actually buying
Build · August 22, 2026 · 1 publisher
- A 27B-parameter agent beat two frontier models at one task, and the task was chosen carefully
Product · August 22, 2026 · 1 publisher
- Meta's coding agent has two prices: pay 18x more, or let it train on your repository
Invest · August 22, 2026 · 1 publisher
- Physics-only world models cannot predict people, and the fix costs six pipeline stages
Build · August 22, 2026 · 1 publisher
- Claude's prompt cache dies quietly in agent loops: the 20-block lookback nobody configures
Build · August 21, 2026 · 1 publisher
- A Government Switched Off Two Frontier Models. Your Board Will Want The Fallback Plan.
Leadership · August 21, 2026 · 1 publisher
- TrueFoundry open-sources an agent harness and calls managed agents a lock-in play
Build · August 19, 2026 · 2 publishers
- Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design
Build · August 19, 2026 · 2 publishers
- Anthropic's protein binders got tested by outside labs. The benchmark is still Anthropic's.
Product · August 19, 2026 · 1 publisher
- Adronite's Codistry makes token count, not context window, the axis of competition
Product · August 19, 2026 · 2 publishers