Claude Code's /cost command now counts cache misses and, when it can, names a likely cause for the last one. Developers who leave sessions open over a break can now see how much usage a cold cache costs them.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
Mistral has added Z.ai's GLM 5.3 to Vibe, its terminal coding agent, and moved the smart-approve classifier onto a fast Mistral model whatever model is active. Teams that pick GLM 5.3 to write code still get a Mistral model deciding when a person must approve an action.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives50
- Confidence45
OpenAI's Codex CLI refresh, builds 0.156.0 to 0.160.0, pitches /fork as a separate checkout, but its command reference says /fork clones only the chat. Two agents on one repo need a worktree session, started from the new agent command center.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence45
Bitdefender's free AI Guardian beta vets each Claude Code and OpenClaw action on macOS, citing tests that manipulated agents in over a third of cases. That rate is Bitdefender's own, so the case for checking every agent action rests on the company offering the check.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+30
- Incentives75
- Confidence40
Anthropic's docs say Claude Code mods run in their own sandbox, yet can read your files, start processes and make network requests. That sandbox only routes access through one $ API, so the trust decision happens at install.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence60
Anji Xu's open-source Claude Statuspane puts context use, five-hour and seven-day rate limits and session cost in a card above the Claude Code prompt. The readout lives on that one developer's screen, so whoever answers for a team's total agent spend still needs records kept somewhere central.
Reality
- Evidence45
- Adoption5
- Hype gap0
- Incentives
- Insufficient
- Confidence50
David Heinemeier Hansson's agent-written Rust port of Campfire served 36,120 room-page requests a second to Rails' 217 in 37signals' own benchmark. A parity harness checks its behavior against Rails, but the numbers are company-run and stop short of production traffic.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives50
- Confidence40
Glow says AI coding agents asked for review screenshots put over 13,000 internal images from 300-plus organizations into public GitHub repositories. Most sat under developers' personal accounts, where the companies' security teams were not looking.
Perspective Coverage
5 publishers
- Builder
- Builder 37%
- Operator
- Operator 53%
- Investor
- Investor 10%
Reality
- Evidence55
- Adoption45
- Hype gap+15
- Incentives70
- Confidence60
Addy Osmani's Opus 5.5 guide for Anthropic says to delete 'think carefully' lines and give each task a finish line and one stop condition. Its sturdier advice covers long Claude Code runs, where a CLAUDE.md rule tells the model when to keep going and when to stop.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence35
Anthropic's CI job volume grew 25x in six months as coding agents raised pipeline load, The New Stack reports. Faster runners and test selection cut the cost of each run, yet repo tests still mock the service seams where distributed systems break.
Reality
- Evidence45
- Adoption60
- Hype gap+20
- Incentives50
- Confidence45
AWS's project spending limits pause a project at its cap and permanently delete its data after 90 days paused without action. The cap bounds what a runaway agent experiment can bill, so someone has to own recovery and keep independent backups.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence55
LDraw Nova, Carlos Antelo's open-source tool, had Claude Opus 5.5 design a 2,175-piece LEGO garden by writing a Python program that emits the CAD file. Nova checks parts for collisions but not stability, so whether any of its designs would stand up in real bricks is still untested.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence40
Anthropic's Claude Code mods, on by default from 2.1.287, let JavaScript or TypeScript code rewrite prompts, block tool calls and approve permission requests. For teams, the governance question moves from what the agent is told to which code it runs.
Perspective Coverage
5 publishers
- Builder
- Builder 59%
- Operator
- Operator 36%
- Investor
- Investor 5%
Reality
- Evidence78
- Adoption15
- Hype gap+15
- Incentives60
- Confidence72
AWS added a per-project spend limit on September 16, 2026 that pauses service at the monthly cap, following Google Cloud's July launch of Spend Caps. A cap that checks each request stops at once, while one built on lagging billing data keeps charging until a function fires.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence30
Shopify is moving every mobile app off React Native to Swift and Kotlin, with its Shop app in stores 12 weeks after a one-week agent-built prototype. The shipped app came out of a gated, screen-by-screen review process that a team copying Shopify would have to build first.
Reality
- Evidence40
- Adoption30
- Hype gap+10
- Incentives
- Insufficient
- Confidence45
Kiro Workflows define multi-agent coding runs as JSON or YAML graphs of five node types that the Kiro Runtime executes in the background. The runtime enforces loops and joins, but inputs reach agents as untyped text, so a deploy flag is only as safe as a model's reading of it.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence40
OpenAI has fixed a Critical Codex flaw in which a semicolon in a branch name leaked the agent's GitHub OAuth token on all four Codex surfaces. How far one leaked token could reach was set by its scope, which is chosen by whoever provisions the agent.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives45
- Confidence60
Anthropic's Claude Sonnet 5.5 beats Opus 5.5 at coding for half the per-token price, on its own tests and on Artificial Analysis's. At max effort it writes 60% more tokens per task, so moving coding work down a tier saves nearer a fifth than a half.
Perspective Coverage
4 publishers
- Builder
- Builder 39%
- Operator
- Operator 36%
- Investor
- Investor 25%
Reality
- Evidence60
- Adoption30
- Hype gap+25
- Incentives65
- Confidence58
Bouras, Dai and Mechtaev found a static denylist let 46 of 75 prompt injections execute in a coding agent, against 3 under preflight-scoped capabilities. A same-day Google report of malware stealing OIDC tokens from GitHub Actions runners puts the outer limit on an agent in the CI job's permissions.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
Earlier coverage
- Claude Code Pro and Max subscribers have until October 7 to claim up to $250 in cloud-session credit
Build · October 3, 2026 · 1 publisher
- 37signals now treats hand-written code as a bug to trace back to its agents
Build · October 3, 2026 · 1 publisher
- Cloudflare's Issues for Workers needs a triage step before errors reach a coding agent
Build · October 3, 2026 · 1 publisher
- AgSpec nearly doubles speculative draft length by indexing code the way coding agents write it
Build · October 2, 2026 · 1 publisher
- Steve Weis factors RSA-896 in ten days with Claude and idle GPUs
Build · October 2, 2026 · 1 publisher
- A Claude Code permission rule now enforces the commit ban that 18 prompt files only requested
Build · October 2, 2026 · 1 publisher
- Cloudflare's cloudflared 2026.9.3 adds optional one-time-PIN email gating for Quick Tunnel links
Build · October 2, 2026 · 1 publisher
- Plain Claude Code drew AWS diagrams as well as official MCP servers once the prompt was tuned
Build · October 2, 2026 · 1 publisher
- Agent work needs separate fields for acceptance, merge and deployment, a dev.to tutorial argues
Build · October 2, 2026 · 1 publisher
- PreToolUse hooks now enforce the rules one developer's coding agent kept skipping
Build · October 2, 2026 · 1 publisher
- Coding models hard-coded answers to example tests they had flagged as wrong
Build · October 2, 2026 · 1 publisher
- Pi 1.0 runs MCP tools inside a sandbox so their output stays out of context
Build · October 2, 2026 · 1 publisher
- CUDA engineers now supervise the AI that writes their GPU kernels
Leadership · October 2, 2026 · 1 publisher
- Griffiths Waite would move shared UI spending from component code to design tokens and tests
Build · October 2, 2026 · 1 publisher
- Microsoft publishes token and cache telemetry from 301,026 Copilot agent sessions
Build · October 1, 2026 · 1 publisher
- Factory's CEO accuses an advisor turned Cognition executive of sharing confidential information
Product · September 30, 2026 · 2 publishers
- OpenAI moves Codex fully into the cloud in the year its agents bypassed their sandboxes
Product · October 1, 2026 · 2 publishers
- IBM's self-hosted Bob coding agent runs air-gapped on Nvidia and Poolside models
Product · October 1, 2026 · 2 publishers
- Claude Code inferred a green build from silence when an approval prompt blocked its exit-code check
Build · October 1, 2026 · 1 publisher
- NVIDIA writes BlueField's API contracts into SKILL.md files for coding agents
Build · October 1, 2026 · 1 publisher
- Khosla calls portfolio company Factory 'second tier' in a fight over a board observer's move to Cognition
Leadership · October 1, 2026 · 1 publisher
- Pi adds MCP through a JavaScript sandbox that skips tool-schema preloading
Build · September 30, 2026 · 1 publisher
- Reading the clock before SQLite's write lock made two AI-built queues issue expired leases
Build · October 1, 2026 · 1 publisher
- DHH says 37signals treats hand-written code like a bug in Sentry
Build · October 1, 2026 · 1 publisher
- JetBrains Air pulls coding agents out of AI chat and into parallel editor sessions
Build · October 1, 2026 · 1 publisher
- Breaks longer than the cache TTL added about 15% to one developer's Claude Code input usage
Build · October 1, 2026 · 1 publisher
- Coding agents routed more than 13,000 private screenshots through public GitHub repos
Build · October 1, 2026 · 3 publishers
- Chrome switches to two-week releases from version 153 on desktop, Android and iOS
Build · October 1, 2026 · 1 publisher
- Snorkel AI's LibraryDesignBench grades agent-written libraries by the code other agents write with them
Build · September 30, 2026 · 1 publisher
- Factory's CEO says a board observer leaked to Cognition before becoming its revenue chief
Build · September 30, 2026 · 2 publishers
- Factory says its board observer spent weeks talking to Cognition before joining it as revenue chief
Leadership · September 30, 2026 · 1 publisher
- Kong AI Gateway hides the rollback tool from Muse Code's investigator by filtering tools/list per identity
Build · September 30, 2026 · 1 publisher
- Cloudflare's Workers Issues hands grouped production errors to Claude Code, Cursor or Devin
Build · September 30, 2026 · 1 publisher
- Git 2.56 ships its conflict-marker check as a flag that scripts and agents must opt into
Product · September 30, 2026 · 1 publisher
- Relipa's coding-agent pipeline sets the tests before the agent writes any code
Build · September 30, 2026 · 1 publisher
- Sign in with ChatGPT now counts Amp and Devin agent runs against Plus and Pro plan limits
Build · September 29, 2026 · 2 publishers
- Claude Code 2.1.278 keeps quotes in $ARGUMENTS and strips them from positional arguments
Build · September 29, 2026 · 1 publisher
- OpenAI's Codex now saves a project's setup so every cloud task starts with dependencies installed
Product · September 29, 2026 · 2 publishers
- A 755-line AGENTS.md moved one of 26 assertions in a controlled agent test
Build · September 29, 2026 · 1 publisher
- One developer's step-level audit finds 97 of 200 Claude Code steps fit a local model
Build · September 29, 2026 · 1 publisher