Hermes Agent v0.21.2 fills passwords from 1Password, Bitwarden or its own vault into web pages without the model ever seeing them. Every card fill waits for a person, so the agent cannot charge a stored card on its own, cron jobs included.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
Lightpanda shipped 1.0 of its open-source browser on October 2nd, passing 1,739,845 web-platform subtests without rendering pages. One encoding suite makes up about two-thirds of that count, so teams should judge it by whether their jobs need to see the page.
Reality
- Evidence45
- Adoption35
- Hype gap+25
- Incentives60
- Confidence50
Google's Cloud Run instances, priced from $5.70 a month, kept an open-source AI agent running with its memory intact in a developer's Preview test. For an agent that mostly waits, it trades a VM's patching for limits the code must be written around.
Reality
- Evidence40
- Adoption8
- Hype gap+15
- Incentives40
- Confidence38
Polyglot's developer ran six coding agents 30 times on each of seven local models, and three never made a tool call on models that write calls as text. The author's own error bars say 30 runs can sort agents into tiers but cannot rank two agents inside one.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives70
- Confidence55
ThreatDown says the Carbonato worm breaks into Docker daemons open on port 2375 and installs the open-source Hermes AI agent, then rescans nearby networks every five minutes to spread. An operator drives each infected host over Telegram.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+10
- Incentives35
- Confidence50
CARBONATO, a Docker botnet running since October 2024, hijacks hosts on open port 2375 and steals keys from 14 AI providers to fund its own LLM gateway. ThreatDown found the crew's own container registry exposed, handing defenders 4.3 GB of its toolchain.
Perspective Coverage
3 publishers
- Builder
- Builder 33%
- Operator
- Operator 57%
- Investor
- Investor 10%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+20
- Incentives35
- Confidence60
OpenClaw's Telegram enabled flag stopped inbound polling while cron-driven sends kept firing hourly through three stale bot accounts, one operator reports. The dashboard showed the channel off, and the sends ended only when the bot tokens left the secrets directory.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
Manifold Security reported eight instances across seven command-line agents. The command runs as the user, outside the sandbox, with no approval prompt, and four were still firing when the researchers retested on September 1.
Publishers:manifold.security · thehackernews.com Reality
- Evidence72
- Adoption60
- Hype gap+15
- Incentives55
- Confidence70
The new Model Context Protocol endpoint reaches every device and the whole event history in Google Home, but only for US subscribers on the paid Premium Advanced tier, and Google is shipping it as early access.
Perspective Coverage
12 publishers
- Builder
- Builder 35%
- Operator
- Operator 50%
- Investor
- Investor 15%
Reality
- Evidence70
- Adoption10
- Hype gap+35
- Incentives60
- Confidence75
Gambit Security says one operator used three AI agent tools to steal over 600,000 cards, spending an estimated $12,000 to $18,000. One retailer's scheduled job reinstalled a skimmer after a deploy removed it, so recovery tests have to cover every layer the agents touched.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives70
- Confidence45
A health dashboard update from AWS reports customer data lost for good in the Middle East. The account carrying its wording to me is a Thai-language post that a model drafted and a human editor checked.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence60
Gambit says a single actor pointed the Strix, Cairn and Hermes agent frameworks at online retailers from July onward, taking more than 600,000 valid card records from two victims and leaving skimmers on 119 sites.
Reality
- Evidence58
- Adoption62
- Hype gap+14
- Incentives62
- Confidence60
The company known for a pocket gadget now ships a cloud agent that installs on up to five Windows, macOS or Linux machines and runs on whatever model subscription the user already pays for. Four rival agents got there first.
Reality
- Evidence32
- Adoption12
- Hype gap+28
- Incentives72
- Confidence44
Inception's diffusion model is the fastest endpoint in the cheap tier on both figures, but OpenRouter's median sits at less than half the vendor number and only about 15 percent above Gemini 3.5 Flash-Lite's measured 382 tok/s.
Reality
- Evidence45
- Adoption58
- Hype gap+28
- Incentives62
- Confidence42
OpenRouter has dropped the stealth listing. The model is Pareto 26.9, from unbiased.ai, at $2.50 per million input tokens before a 10 October launch. That is double what 26.8 charged.
Reality
- Evidence64
- Adoption42
- Hype gap+45
- Incentives68
- Confidence55
A dev.to comparison scores MCP tool design at 13 website builders and content systems against eight criteria. The two designs it walks through in full both take deletion out of the flat tool list and put it behind its own gate.
Reality
- Evidence46
- Adoption40
- Hype gap+18
- Incentives28
- Confidence45
A Thai-language walkthrough estimates 15 to 20 MCP calls to build a ceramics studio site on WordPress.com. Counting each documented operation, including one status update per item to publish, gets to 22.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap−10
- Incentives28
- Confidence42
A dev.to post pulled every figure for 12 popular Claude Skills straight from the GitHub API on 19 September 2026, then set the result against what three skill catalogs advertise for the same repository.
Reality
- Evidence45
- Adoption50
- Hype gap−10
- Incentives25
- Confidence45
GitSpawn, as summarised in a dev.to write-up of Cloud Security Alliance findings, uses a Git performance setting to run attacker code when a coding agent inspects a project it just opened. Seven agents are named.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+18
- Incentives28
- Confidence40
Nous Research spent about $19,300 on 1,393 subagents to take 34.4 percent out of Hermes's non-test Python in 19 hours. Test line count moved 0.06 percent, so the guardrail was the suite the team already had.
Reality
- Evidence38
- Adoption24
- Hype gap+32
- Incentives74
- Confidence46
Earlier coverage
- Google hands five outside coding agents access to its Home devices
Product · September 17, 2026 · 1 publisher
- Four conditions gate every emergency container this watchdog is allowed to start
Build · September 16, 2026 · 1 publisher
- A unanimous verifier panel gates every proof in Nvidia's IMO gold recipe
Build · September 13, 2026 · 1 publisher
- Opening an untrusted repo made seven coding agents run the program its Git config named
Build · September 13, 2026 · 1 publisher
- Nine of 428 audited LLM routers injected code into the tool calls they relayed
Build · September 12, 2026 · 1 publisher
- Requests to deepseek-v4-pro start returning V4.1-Flash on 14 September at 04:00 UTC
Build · September 11, 2026 · 1 publisher
- GLM-5.3-Flash buys seven retries for the price of one Kimi K3 call
Build · September 11, 2026 · 1 publisher
- A wiki that accepted GET as an edit gave read-only agents 18,000 writes
Build · September 11, 2026 · 1 publisher
- A NetScaler web shell survives the patch that closes CVE-2026-8452
Build · August 29, 2026 · 1 publisher
- Four agent runtimes, four blast radii: the teammate interface converged, isolation did not
Build · August 24, 2026 · 1 publisher
- Superpowers makes spec-driven work a precondition, then ships it to twelve harnesses
Build · August 24, 2026 · 1 publisher
- Nous wants you to own the harness, which means you own the patching too
Build · August 23, 2026 · 1 publisher
- Long Horizon: Google open-sources the agent bugs that never threw an error
Build · August 22, 2026 · 1 publisher
- Solar Pro 4 turns model routing into a procurement decision, not a research one
Build · August 20, 2026 · 1 publisher
- Six pragmas and a context manager: the vector store that fits in 2GB of RAM
Build · August 19, 2026 · 1 publisher
- Eight agents, four days, 1,395 files: the AI intrusion campaign that mostly ran itself
Security · August 18, 2026 · 2 publishers
- Seven agentic AI incidents, one front door: the identity metadata you publish on purpose
Security · August 14, 2026 · 2 publishers
- Wiring, not headcount: same agent task swung from 70% worse to 81% better on topology alone
Build · August 15, 2026 · 1 publisher