A 73-page preprint evolved instructions that jumped between coding agents and wrote themselves into the file that becomes the next system prompt. A short warning nearly stopped transmission.
Perspective Coverage
4 publishers
- Builder
- Builder 52%
- Operator
- Operator 39%
- Investor
- Investor 9%
Reality
- Evidence68
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence65
The multipliers in circulation for how far AI seats are underpriced run from five times to twenty, and all of them rest on an unpublished per-seat token count. The API rate card is the part you can check.
Reality
- Evidence40
- Adoption35
- Hype gap+35
- Incentives58
- Confidence38
Anthropic says Claude Mythos Preview found and exploited previously unknown flaws in every major operating system and web browser during a month of testing. Its disclosure process keeps the rest unnamed until patches ship.
Publishers:red.anthropic.com
Reality
- Evidence36
- Adoption20
- Hype gap+35
- Incentives78
- Confidence55
LangChain clocked TypeSafe AI's Jev at 0.44 seconds and $0.00035 per call against three LLM judges on the same eval set. Whether that price transfers depends on how much structure your traces already have.
Publishers:langchain.com
Reality
- Evidence45
- Adoption15
- Hype gap+18
- Incentives70
- Confidence55
Strands Evals 1.0 added chaos testing and red teaming in June. One AWS builder pointed both at a fixture-fed copy of his cost-reporting agent, because Cost Explorer charges a cent per API request.
Reality
- Evidence40
- Adoption15
- Hype gap+30
- Incentives55
- Confidence35
Antigravity's free tier includes Gemini 3.1 Pro and Claude Opus 4.6 with unlimited tab completions. XDA reports Opus burns credits about four times faster than Gemini, and Google moved the free tier onto credits in March 2026.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence40
The Unified Harness Protocol specifies how an application starts a task on an agent runtime, follows it, cancels it and collects the files, borrowing the shape of OpenAI's Responses API so existing streaming clients need no changes.
Publishers:unifiedharnessprotocol.org
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+35
- Incentives62
- Confidence52
Two models read the same tool description and disagreed about one field name. The fix moved the shape into the JSON Schema for the nine pattern kinds that account for 85 percent of emissions, and left the rest loose.
Reality
- Evidence45
- Adoption20
- Hype gap−10
- Incentives25
- Confidence55
Newer models ship with advice to keep the guidance file under 200 lines. The analyst who wrote those 1,042 lines argues each one logs context the model lacked, and that the real debt is a mechanical rule left sitting in prose.
Reality
- Evidence42
- Adoption22
- Hype gap−12
- Incentives38
- Confidence55
The Kotlin Benchmark grades agents on 105 verified repository tasks. Its token column shows setups that solve within a few tasks of each other burning between 66,000 and 777,000 tokens per fix.
Reality
- Evidence62
- Adoption30
- Hype gap+22
- Incentives60
- Confidence58
AWS puts prompt caching's saving at up to 90 percent on cache hits; its own ten-question example nets about 75 percent, and only while every request lands inside the five-minute default TTL.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+18
- Incentives80
- Confidence58
The harness lost its hidden system prompt, 43% of its builtin tool descriptions and its todo list middleware. LangChain's own footnote says reward confidence intervals span zero for every model tested, so the evals settle the token saving more firmly than the quality.
Publishers:langchain.com
Reality
- Evidence58
- Adoption30
- Hype gap+18
- Incentives82
- Confidence46
AWS shows three agents sharing one AgentCore container across two hosting paths. The orchestration framework absorbs the split; the OpenTelemetry instrumentation does not.
Reality
- Evidence74
- Adoption
- Insufficient
- Hype gap+8
- Incentives86
- Confidence66
A no-code support agent called an endpoint it invented, was refused twice, and still told the customer it could see her charges. Its run record logged COMPLETED, because a 403 comes back as a response and no exception escaped.
Reality
- Evidence58
- Adoption10
- Hype gap−6
- Incentives68
- Confidence45
A study of seven agent harnesses reports 770 confirmed passes in 1,000 runs of a plugin-update attack, and no run was blocked by the model. The harness dispatches the hook, so the model has nothing to refuse.
Reality
- Evidence57
- Adoption
- Insufficient
- Hype gap+14
- Incentives56
- Confidence53
Anthropic reports that users accept 93% of Claude Code permission prompts. Its replacement is a classifier that judges each action, and switching it on revokes the blanket allow rules teams wrote for convenience.
Reality
- Evidence44
- Adoption20
- Hype gap+16
- Incentives76
- Confidence52
Pengcheng Xu's AgentConnect comparison finds semantic navigation beats grep only where text search is noisy; on clean repositories it added 16 to 19 percent to token cost for little or no accuracy gain.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+12
- Incentives72
- Confidence52
Vedere Labs did get Claude Code to move a working exploit onto a second WAGO controller, and the price of doing it puts a measurable ceiling under the pitch that AI now writes OT exploits at machine speed.
Reality
- Evidence50
- Adoption12
- Hype gap−10
- Incentives55
- Confidence55
Forescout's Vedere Labs ported a known WAGO exploit to a new model with Claude and Ghidra, tracking every dollar and human correction. A single bad write to flash memory bricked the target PLC for good.
Reality
- Evidence48
- Adoption12
- Hype gap−12
- Incentives65
- Confidence40
Three separate evaluations, collected in one dev.to post, keep landing on the same defect, which is that the reviewer runs after the side effect. Moving the deny to the tool call is a runtime change, not a prompt change.
Reality
- Evidence58
- Adoption18
- Hype gap+14
- Incentives30
- Confidence52
Earlier coverage
- Copilot's credit meter moves the cost decision into the model dropdown
Build · August 31, 2026 · 1 publisher
- Meta wants up to $199.99 a month for an agent, and it is selling a meter
Product · August 26, 2026 · 1 publisher
- Claude Code tells you the model, not the culprit: 106 lines of shell to name the Skill
Build · August 25, 2026 · 1 publisher
- An AI ops agent's real permissions design is two Istio policies and one ClusterRole
Build · August 24, 2026 · 1 publisher
- Tier the models; the validation boundary is the thing you are actually buying
Build · August 22, 2026 · 1 publisher
- A note checker with no accuracy figure, and the labelled dataset it borrowed to show its misses
Build · August 21, 2026 · 1 publisher
- A goal that writes itself into SOUL.md: agent memory is now an attack surface
Build · August 19, 2026 · 1 publisher
- That 34.3% TREM2 Hit Rate Belongs To Six Agents, Not To Claude
Build · August 18, 2026 · 1 publisher
- A paragraph beat the agent "mind virus": reading the Anthropic-EPFL preprint as a defensive win
Security · August 18, 2026 · 1 publisher
- The model is now choosing the extortion targets, not just writing the malware
Security · August 18, 2026 · 1 publisher
- A correct anomaly detector and a $30,141.33 Bedrock bill that never tripped it
Build · August 16, 2026 · 1 publisher
- The AI store manager did not fire anyone until humans told it to read its own policy
Product · August 15, 2026 · 1 publisher
- Three Claude agents, one task, and a malware turf war: the multi-agent bill arrives
Invest · August 14, 2026 · 1 publisher