Google researchers' Dream-RSI cut Gemini calls on a Lasso solver task from 550 to 317 by rewriting a Python search policy, with every model frozen. Teams running scored code search can test it as a call-budget saving on their own tasks.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+60
- Incentives
- Insufficient
- Confidence40
Air Security found four coding agents skipped checking plugin code against its pinned commit SHA, a gap one test plugin used to reach 26,000 agents. Fixes are uneven across vendors, so a team's exposure depends on which agent it runs and where its plugin repos are hosted.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence35
The MIT-licensed dsh runtime bundles sessions, tool calls, permissioning and a local web UI, with model adapters as plugins covering Anthropic, OpenAI, Bedrock, Vertex and Azure. It is still a preview.
Reality
- Evidence62
- Adoption35
- Hype gap+20
- Incentives50
- Confidence58
Anthropic's open SKILL.md format is read by roughly forty agent tools, among them Codex, Cursor, Copilot and Gemini CLI. Each installed skill costs about 100 tokens a session until a task matches its description and the full workflow loads.
Reality
- Evidence45
- Adoption45
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
GKE Agent Sandbox is generally available and publishes provisioning numbers a team can plan against. The scheduling and durability layers stacked on top of it rest on one public demo, and the deprecation argument lands there.
Reality
- Evidence48
- Adoption14
- Hype gap+58
- Incentives50
- Confidence55
Plugin4Shell lets attacker code replace SHA-pinned plugins in Claude Code, Codex, Gemini CLI and Copilot through default auto-update, with no prompt. A pin only protects you if the agent checks fetched code against the hash, so agent plugins need the same review as any other dependency.
Reality
- Evidence25
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence25
Claude Code started reading AGENTS.md on September 18, but only where no CLAUDE.md sits in the working directory or any parent above it. Monorepos need a stub file per package or a change to the instructionFiles setting.
Reality
- Evidence55
- Adoption40
- Hype gap+10
- Incentives35
- Confidence55
Oren Yomtov of Accomplish found two ways out of the Codex sandbox, both silent and neither stopped by an approval prompt. OpenAI patched them within eight days, and other researchers have found the same design in rival agents.
Reality
- Evidence68
- Adoption45
- Hype gap−5
- Incentives55
- Confidence62
Air Security's Plugin4Shell lets a plugin be swapped during a background refresh while the agent reports the audited commit. Anthropic and OpenAI have patched, Copilot has no fix, and Google is retiring Gemini CLI.
Reality
- Evidence60
- Adoption65
- Hype gap+20
- Incentives70
- Confidence58
The repo advertised a token-cost dashboard that read agent session logs locally. According to a dev.to post, line 12 of its CLI called into code that pulled a module from a bare IP and exec'd it in RAM, and a later commit moved that payload inline.
Reality
- Evidence62
- Adoption30
- Hype gap+12
- Incentives58
- Confidence55
Researchers at AIR found that four AI coding agents accepted a Git checkout they never verified, and because no marketplace can close the gap, the repair now travels through agent version numbers inside each enterprise.
Reality
- Evidence62
- Adoption38
- Hype gap+18
- Incentives60
- Confidence55
Oren Yomtov of Accomplish AI found two escapes from the Codex sandbox. The worse one ran unsandboxed commands from read-only mode through a helper tool that Codex Desktop writes into the global config at install.
Reality
- Evidence72
- Adoption45
- Hype gap+10
- Incentives55
- Confidence58
Anthropic shipped the AGENTS.md fallback in Claude Code 2.1.277 as a flag-gated built-in mod. Two machines on the same pinned version can load different project instructions. The same week's digest also logs a reverted deny-rule fix.
Perspective Coverage
4 publishers
- Builder
- Builder 52%
- Operator
- Operator 39%
- Investor
- Investor 9%
Reality
- Evidence78
- Adoption62
- Hype gap−6
- Incentives58
- Confidence74
The Unified Harness Protocol specifies how an application starts a task on an agent runtime, follows it, cancels it and collects the files, borrowing the shape of OpenAI's Responses API so existing streaming clients need no changes.
Publishers:unifiedharnessprotocol.org
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+35
- Incentives62
- Confidence52
In AIR's Plugin4Shell research, a branch named after the hash you pinned wins the checkout in Claude Code, Codex, Copilot and Gemini CLI. The check that would catch it is one line of git, run on your disk.
Reality
- Evidence34
- Adoption27
- Hype gap+45
- Incentives58
- Confidence33
Check Point's July-August digest has evaluation models from OpenAI, Anthropic and Meta reaching production systems, and only the OpenAI model got there by finding a bug. The other two environments were left reachable.
Reality
- Evidence30
- Adoption38
- Hype gap+38
- Incentives76
- Confidence30
Paper2Agent tests every function it extracts against the paper's own published output before shipping it. Its 26 failures out of 100 papers also put a rough ceiling on how many computational biology repositories still run.
Reality
- Evidence58
- Adoption22
- Hype gap+8
- Incentives65
- Confidence55
Google's server for Google Analytics is a read-only Python package you run locally through pipx, and it has been on GitHub since mid-2025. Anthropic only handed the protocol to the Linux Foundation that December.
Reality
- Evidence48
- Adoption55
- Hype gap+14
- Incentives38
- Confidence50
The Kotlin Benchmark grades agents on 105 verified repository tasks. Its token column shows setups that solve within a few tasks of each other burning between 66,000 and 777,000 tokens per fix.
Reality
- Evidence62
- Adoption30
- Hype gap+22
- Incentives60
- Confidence58
Earlier coverage
- Exaforce ties every AI agent it discovers back to the employee whose permissions it borrows
Product · September 15, 2026 · 1 publisher
- Claude Code's deny rules on /tmp and /etc missed paths given by their real location
Build · September 12, 2026 · 1 publisher
- Prompt injection embedded in malware turns an LLM scanner's refusal into a free pass
Security · September 11, 2026 · 1 publisher
- Anthropic and GitHub have moved AI costs from the seat to the meter
Leadership · September 3, 2026 · 1 publisher
- Team-scoped gateway keys tag org, team and user identity for AI spend dashboard
Build · August 30, 2026 · 1 publisher
- Hidden text in a public bug report walked out with Editor rights on a Google Cloud project
Product · August 28, 2026 · 1 publisher
- Twenty-three security checks, zero coverage: AI coding agents as build-pipeline attack surface
Build · August 25, 2026 · 1 publisher
- An agent guard that runs on your laptop, and cannot tell you whether anyone keeps it on
Security · August 24, 2026 · 1 publisher
- Superpowers makes spec-driven work a precondition, then ships it to twelve harnesses
Build · August 24, 2026 · 1 publisher
- The hard part of running five coding agents is not the model, it is the process tree
Build · August 23, 2026 · 1 publisher
- Parallel coding agents on Windows break at the home directory, not the launcher
Build · August 23, 2026 · 1 publisher
- One agent, 119 blog heroes, and the scaffolding that made them shippable
Build · August 22, 2026 · 1 publisher
- The three bugs that decide whether an agent office survives the night
Build · August 22, 2026 · 1 publisher
- The agent did not fail, the client did: 90 logged MCP trials and a validator that ate the calls
Build · August 21, 2026 · 1 publisher
- Claude Code's agent-team panes need tmux, and Anthropic says Windows Terminal is out
Build · August 20, 2026 · 1 publisher
- The trust prompt is the attack surface: cloned repos can spawn MCP servers with your privileges
Science · August 19, 2026 · 1 publisher
- OX Security says MCP command execution is a design choice, so server owners own the risk
Science · August 19, 2026 · 1 publisher
- Before you spend quota on an agent skill, make it pass an eval harness
Build · August 18, 2026 · 1 publisher
- The linter that passed everyone who ignored it and warned everyone who complied
Build · August 17, 2026 · 1 publisher
- Claude Code's new default is a confession: the approval prompt was never a control
Build · August 16, 2026 · 1 publisher
- 65,000 pulls a day, one author: the AI coding stack's unpriced dependency
Invest · August 15, 2026 · 1 publisher