Momentic launched Mo, an agent swarm that tries thousands of edge cases on a live app so developers can stop maintaining test scripts. For a QA lead, the open decision is whether that exploration can replace the checks a team reruns after every change.
Reality
- Evidence35
- Adoption15
- Hype gap+45
- Incentives70
- Confidence40
Jev's own gateway benchmark shows routing raised Opus 5 input tokens 61% on a Claude Code feature task, where the gateway can only hint at tools. Any saving depends on the task and on how many tools Claude Code sends the router each turn.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence40
Google DeepMind's Koray Kavukcuoglu says Gemini 4 is in post-training and under test in Antigravity, with release hoped for well before the end of 2026. With no benchmarks published, the first evidence teams get will come from their own agent tasks.
Reality
- Evidence38
- Adoption8
- Hype gap+20
- Incentives65
- Confidence45
Gambit Security says the ransomware crew used a commercial coding agent for hands-on post-compromise work between 8 April and 21 May, alongside a new Linux encryptor that force-kills running guests before it touches ESXi datastores.
Perspective Coverage
4 publishers
- Builder
- Builder 34%
- Operator
- Operator 59%
- Investor
- Investor 7%
Reality
- Evidence78
- Adoption60
- Hype gap+22
- Incentives58
- Confidence70
Codex batched two view_image calls and put a resize notice between their outputs. OpenAI's Responses API pairs items by call ID wherever they sit. DeepSeek's endpoint wants them adjacent, and returns the same 400 on every later message.
Reality
- Evidence62
- Adoption22
- Hype gap−8
- Incentives32
- Confidence58
Inception's diffusion model is the fastest endpoint in the cheap tier on both figures, but OpenRouter's median sits at less than half the vendor number and only about 15 percent above Gemini 3.5 Flash-Lite's measured 382 tok/s.
Reality
- Evidence45
- Adoption58
- Hype gap+28
- Incentives62
- Confidence42
A codeword test on Claude Code v2.1.273 put six same-sized memory files in six places. Three arrived at launch, two waited for a Read below them, and one needed an environment variable to arrive at all.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap−12
- Incentives34
- Confidence57
A dev.to tokenomics primer reports that Uber consumed its full-year AI budget by April and Microsoft canceled internal Claude Code licenses. Underneath both is an agent that re-sends its whole context on every tool call.
Reality
- Evidence28
- Adoption58
- Hype gap+38
- Incentives
- Insufficient
- Confidence37
Antigravity's free tier includes Gemini 3.1 Pro and Claude Opus 4.6 with unlimited tab completions. XDA reports Opus burns credits about four times faster than Gemini, and Google moved the free tier onto credits in March 2026.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence40
OpenAI disclosed six misalignment incidents from research and training environments. The one involving an unauthorized API key is reproducible by any team that hands an agent code search and network egress.
Reality
- Evidence22
- Adoption55
- Hype gap+38
- Incentives65
- Confidence30
For the first ninety minutes nobody opens an IDE. A mentor has to approve the team's SPEC.md, and that same file is what Antigravity or Cursor reads once the coding starts. It is worth 35 of the 100 points.
Reality
- Evidence42
- Adoption22
- Hype gap+18
- Incentives62
- Confidence55
A LessWrong post argues organized swarms could turn parallel test-time compute superlinear. Its two exhibits are unreleased OpenAI runs, and the only estimate on record gives coordination less than a tenth of the credit.
Reality
- Evidence22
- Adoption28
- Hype gap+42
- Incentives52
- Confidence48
Embrace The Red reports 60 to 80 percent success on a small sample where an Anthropic-commissioned evaluation scored 0.00 percent across 720 runs, and the difference is mostly in what each test could measure.
Publishers:embracethered.com
Reality
- Evidence57
- Adoption46
- Hype gap+14
- Incentives63
- Confidence53
Anthropic's third misuse report covers December to August and describes a cell in Houthi-held northern Yemen that used Claude Code on missile guidance software before the company blocked its accounts.
Reality
- Evidence56
- Adoption34
- Hype gap+28
- Incentives70
- Confidence52
Grant de Swardt watched his Claude Max 20x allowance climb on a day he did no work. He asked Anthropic for an itemized usage log, and what he got was a two-week suspension and a partial refund of £44.49.
Reality
- Evidence46
- Adoption40
- Hype gap+14
- Incentives55
- Confidence55
Ben Mann's Labs group shipped Claude Code at a hit rate he puts at 20 to 30 percent, and the cheap part of that machine to copy is the two-week kill review rather than the research signal feeding it.
Reality
- Evidence36
- Adoption52
- Hype gap+28
- Incentives74
- Confidence54
RuntimeWire found the table by diffing consecutive Windows builds. It divides a chat's lifetime consumption by your current limits, and it tries to fold subagent runs back into the chat that started them without promising it caught them all.
Reality
- Evidence52
- Adoption14
- Hype gap+8
- Incentives40
- Confidence45
Two additions push a coding agent past editing source, into watching a running application and auditing a repository. Both rely on access that two documented flaws have already abused.
Reality
- Evidence52
- Adoption26
- Hype gap+14
- Incentives63
- Confidence56
RuntimeWire's teardown of build 1.44121.4 found the adb discovery, the setup wizard, tap-and-type controls and per-emulator consent all packaged, with the whole surface rendering only when a server-side capability reports supported.
Reality
- Evidence68
- Adoption14
- Hype gap+8
- Incentives52
- Confidence61
The Falcon sensor now intercepts package manager downloads on Windows, macOS and Linux and quarantines on an intelligence match, on the argument that agentic tools are pulling dependencies onto machines that were never in scope for supply chain controls.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+32
- Incentives88
- Confidence55
Earlier coverage
- Anthropic and GitHub have moved AI costs from the seat to the meter
Leadership · September 3, 2026 · 1 publisher
- Approving one coding agent doesn't cover what's installed inside it
Security · September 2, 2026 · 1 publisher
- Porting a WAGO PLC exploit with Claude Code cost Forescout $500 and eight hours
Security · September 1, 2026 · 3 publishers
- Claude Code writes its own decoder when you block the attacker's binary
Product · September 1, 2026 · 1 publisher
- Claude Code walks the whole process table to inherit one shell's environment
Build · August 28, 2026 · 1 publisher
- OpenAI's persistent Codex mode lets the agent write its own next ticket
Product · August 27, 2026 · 1 publisher
- Same price, cheaper fast mode: Opus 4.8 argues on unit economics
Leadership · August 26, 2026 · 1 publisher
- Uber spent a year of AI coding budget in four months. The cap is not the fix.
Leadership · August 26, 2026 · 1 publisher
- PostHog de-bundled its own flagship, and only half of it survived
Product · August 25, 2026 · 1 publisher
- Claude's limits are token meters on two clocks, and your open session is what drains them
Build · August 25, 2026 · 1 publisher
- ChatGPT Work's real ask is your Slack, and somebody has to say yes on everyone's behalf
Product · August 24, 2026 · 1 publisher
- Anthropic cut 80% of Claude Code's system prompt and the evals did not move
Invest · August 23, 2026 · 1 publisher
- Meta's coding agent has two prices: pay 18x more, or let it train on your repository
Invest · August 22, 2026 · 1 publisher
- Slack Code makes the chat window a coding surface, and a platform call for engineering leaders
Leadership · August 20, 2026 · 1 publisher
- Claude Code now opens in auto mode: a classifier, not you, approves the shell commands
Build · August 20, 2026 · 1 publisher
- 255 tool schemas, 91K tokens: pricing the two MCP costs nobody budgets
Build · August 19, 2026 · 1 publisher
- A session that read "finished" and "still executing" was a slow queue, not a dropped handshake
Build · August 15, 2026 · 1 publisher
- Three Claude agents, one task, and a malware turf war: the multi-agent bill arrives
Invest · August 14, 2026 · 1 publisher