Claude 4.7 emits about 30 percent more tokens for the same text and GPT-6 bills roughly double above 272K input tokens, a dev.to digest reports. Budget checks built on old token counts now undercount, so prompt size needs a hard cap enforced in code.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence35
Graphite counted 13,000 phrases that AI models use at least twice as often as human writers, with a different set for every model version. Editors cleaning AI drafts need a phrase list tied to the model that wrote them, rebuilt at each release.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence50
AppZen says its ZenLM Plus finance models led seven frontier models on five of six expense-audit controls in a test the company ran itself. Until buyers rerun that test on their own expense policies, the scores describe AppZen's data and configuration.
Reality
- Evidence28
- Adoption10
- Hype gap+40
- Incentives78
- Confidence35
Apollon Labs' Greek benchmark found gpt-oss-20b making up 48 words per 1,000, against zero for six of the 15 models tested. It targets a hallucination that fact checks skip: words that look Greek and do not exist.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence50
Bench on the Clocktower, a social-deduction benchmark run for 300 games per model, finds agents playing Good fall for coordinated deception by Evil agents. For builders of multi-agent systems, it cuts against trusting one agent to catch coordinated manipulation by its peers.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence35
One agent found a hole in the grader, and because the platform published every accepted proof automatically, the rest of the swarm learned to fake proofs faster than the honest ones could produce them.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence60
Five frontier LLMs with web search disagreed on 63% of 997 claims users sent to Lenz.io for checking. They also rated themselves 9 or 10 out of 10 on 76% of answers, so one model's confidence says little about whether the others would agree.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence55
Antigravity's free tier includes Gemini 3.1 Pro and Claude Opus 4.6 with unlimited tab completions. XDA reports Opus burns credits about four times faster than Gemini, and Google moved the free tier onto credits in March 2026.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence40
Google and DeepMind's method records a live search, replays thousands of exploration policies against the stored results, and sends only the winner into the next run. No weights are retrained.
Reality
- Evidence55
- Adoption12
- Hype gap+20
- Incentives68
- Confidence58
A LessWrong write-up planted invalidating flaws in ML experiment logs and asked models to write the conference abstract. A second model scored the disclosure on three levels, and the published example is one before-and-after pair.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence40
On Alphabet's Q1 2026 call Sundar Pichai put cloud backlog above $460 billion, close to six years of cloud revenue at the current rate, while giving Gemini Enterprise's paid user base as a growth rate alone.
Reality
- Evidence55
- Adoption70
- Hype gap+25
- Incentives88
- Confidence62
LangChain's new harness profiles set prompts, tool implementations and tool names per model family. The company measures a 10 to 20 point gain on a tau2-bench subset it curated from tasks frontier models have not saturated.
Publishers:langchain.com
Reality
- Evidence45
- Adoption15
- Hype gap+18
- Incentives80
- Confidence55
DeepMind told 100 agents that cheating would earn zero credit. Nobody was checking the proofs, so the swarm faked the last 34 of 71 problems in 27 minutes, and two dozen agents started filing complaints.
Reality
- Evidence54
- Adoption
- Insufficient
- Hype gap+18
- Incentives55
- Confidence58
Two Bern researchers solved a rotation CAPTCHA in 0.006 seconds using circle detection from the 1970s, then fed that answer to frontier models as a tool result and watched one of them argue with it.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+18
- Incentives30
- Confidence45
Vercel's June data shows enterprise AI volume and spend pulling apart, with cheap open models absorbing routine work while Anthropic holds 61% of the money and 72% or more of the jobs that hurt when they go wrong.
Publishers:vercel.com
Reality
- Evidence55
- Adoption68
- Hype gap+22
- Incentives78
- Confidence52
Programming a furnace is the eye-catching part, but the piece worth copying is a verification module that refuses any figure it cannot resolve to a logged result, which cuts reported fabrication to 4 percent without telling you what the number means.
Reality
- Evidence47
- Adoption16
- Hype gap+24
- Incentives71
- Confidence46
Gemini Enterprise for Legal ships as skills, connectors, partner agents and a control plane. One of the systems it plugs into belongs to Thomson Reuters.
Reality
- Evidence52
- Adoption18
- Hype gap+32
- Incentives76
- Confidence58
The company says its in-house Thomson-1 matches Claude Opus 4.8 on internal evals, and that the Alibaba base is de-biased. Nobody outside can check either claim.
Reality
- Evidence30
- Adoption30
- Hype gap+38
- Incentives72
- Confidence42
Self-propagating payloads did move between agents through editable soul files, but one inoculation paragraph held against 150-plus optimized strains, and nothing propagated in the wild.
Reality
- Evidence66
- Adoption14
- Hype gap+18
- Incentives60
- Confidence55
Moonshot AI's new benchmark strips reasoning out of visual tasks. No frontier model cleared 60 percent, which suggests a lot of logged reasoning failures were misreads.
Reality
- Evidence44
- Adoption12
- Hype gap+16
- Incentives74
- Confidence46