Google Cloud AI Research has released RRSI, an Apache 2.0 tool whose self-rewriting agent harness lifted Terminal-Bench 2.1 scores from 74.2% to 80.2%. Any team can use it commercially, though on tasks the agent never trained against the reported gain falls to 4.7 points.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence40
LiveReview's Maneshwar says Gemini File Search cost over 50 cents a review run because a reasoning model performed each search. A local index and a cheaper model brought runs down to 4 cents.
Reality
- Evidence45
- Adoption10
- Hype gap+15
- Incentives35
- Confidence40
Researchers found an auto-displayed AI answer cut 'I don't know' responses from 35 percent to 1 percent, though the model was mostly wrong. A 10-cent penalty for wrong answers still left abstention at 7 percent, so review tools need more than an abstain button.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence50
Automated harness evolution keeps whatever edits raise its own benchmark score, so the harness ends up fitted to the eval set. Google Cloud AI Research answers with five regularizers lifted from supervised learning.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+35
- Incentives45
- Confidence30
Google has published a price for frontier-tier access aimed at technical staff, and the figure a manager has to defend is per person per month, on a plan whose usage limits are quoted as multiples of a Pro allowance the company has not spelled out.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+30
- Incentives85
- Confidence62
Arize and Fireworks priced ten models on what a completed command-line task costs, across 2,400 runs. The winner on that metric is an open model with the worst pass rate in the study and the thinnest coverage.
Publishers:arize.com
Reality
- Evidence62
- Adoption18
- Hype gap+14
- Incentives75
- Confidence55
Anthropic charges $20 a month for Cowork's file work and computer control while OpenAI gives folder access away free. That pricing gap means the capability most likely to spread across managed laptops is the one that never generates an invoice.
Reality
- Evidence45
- Adoption30
- Hype gap+12
- Incentives40
- Confidence45
Information agents reach Google AI Pro and Ultra subscribers in summer 2026. The action path matters more than the monitoring, because an agent compares options and may phone your business before anyone at your end sees a customer.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+32
- Incentives55
- Confidence33
Half of those rejections say the same thing, that the proposing model pushed severity past the CVSS evidence it had just cited. That is a real finding about the proposer, and it still says nothing about whether a same-family reviewer would have caught it.
Reality
- Evidence44
- Adoption8
- Hype gap+12
- Incentives48
- Confidence38
JetBrains put two frontier models through the same coding agent, which scored them as a tie on tasks solved even though the runs differed by 47% in steps and 2.25x in dollars, and that gap is what procurement actually pays.
Reality
- Evidence48
- Adoption18
- Hype gap+12
- Incentives68
- Confidence42
An eldercare agent published at a live URL puts its two hard rules in Python instead of prompts, and the demo tests them by calling the send function directly, which is the only version of that claim a reader can check.
Reality
- Evidence32
- Adoption8
- Hype gap+14
- Incentives68
- Confidence46
One escalation loop got the wording right but sent it to the wrong person. No unit test could have caught that, because who gets paged and what stops the paging are both resolved outside the function under test.
Reality
- Evidence46
- Adoption8
- Hype gap−18
- Incentives58
- Confidence47
Google put the same 3.7 Flash behind its API, Vertex, Gemini Enterprise and AI Mode in Search on August 13. The payoff is fewer evaluations to run, not a new capability tier.
Reality
- Evidence42
- Adoption38
- Hype gap+14
- Incentives68
- Confidence36
Arize and Fireworks ran ten models against 40 agent tasks and found the cheapest model per finished job also had the worst pass rate. Coverage, not price, is the binding constraint.
Publishers:arize.com
Reality
- Evidence52
- Adoption20
- Hype gap+22
- Incentives78
- Confidence45
Moonshot AI's new benchmark strips reasoning out of visual tasks. No frontier model cleared 60 percent, which suggests a lot of logged reasoning failures were misreads.
Reality
- Evidence44
- Adoption12
- Hype gap+16
- Incentives74
- Confidence46
A dev.to price walkthrough shows two models swapping places by 17% and 42% on the same list prices. For that pair, the crossover sits at ten input tokens per output token.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap−5
- Incentives70
- Confidence58