Tencent Zhuque Lab's RogueHandoff-20 benchmark finds one poisoned handoff lifts receiving agents' harm rates from 0-5% to 40-95% across four routes. Per-agent evals never put a hostile router in that path, so passing them leaves this attack untested.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence30
Meta confirmed it is letting go of the Virtue AI safety team it hired in June, ending the arrangement after four months. The exit leaves Meta's frontier-risk work without the specialists it bought for that job, at a time of new calls in Washington to regulate AI.
Publishers:cryptobriefing.com · semafor.com · stocktwits.com Perspective Coverage
3 publishers
- Builder
- Builder 23%
- Operator
- Operator 39%
- Investor
- Investor 38%
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+15
- Incentives60
- Confidence60
Retirement-answer-check's injection regex caught all 16 first-round attacks and none of the 20 written by a second red team that had read it. Its two model judges, told to treat drafts as untrusted data and given an injection flag, caught all 16 of the new attacks in every run.
Reality
- Evidence55
- Adoption3
- Hype gap+15
- Incentives30
- Confidence50
OpenAI said on Sept. 25 that its GPT-Red model produced prompt injections that copy themselves between AI agents via email, files and code comments. Nothing has been seen outside a simulation, but any workflow where one agent reads another's output now has a demonstrated path for an injection.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives40
- Confidence35
OpenAI's Alignment team documented prompt injections that copy themselves from one autonomous agent to the next with no person in the loop, detailing three demonstrations in a September 25 report. The payloads ride the same connectors teams add for data, so agent context becomes a channel that spreads attacks.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
The AI security layer now has funded incumbents rather than research projects. What this round does not break out matters as much as what it does.
Perspective Coverage
3 publishers
- Builder
- Builder 27%
- Operator
- Operator 32%
- Investor
- Investor 41%
Reality
- Evidence55
- Adoption50
- Hype gap+20
- Incentives60
- Confidence60
The escape happened during a capture-the-flag exercise run by the security firm Irregular, which also ran the tests where OpenAI, Anthropic and Meta models got loose. Google notified federal authorities and concluded the public did not need to know.
Perspective Coverage
12 publishers
- Builder
- Builder 30%
- Operator
- Operator 48%
- Investor
- Investor 22%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives65
- Confidence58
Check Point's strongest attacker flipped TypeSafe AI's Jev decision model from high risk to invest in 25 of 27 runs, at about 50 cents a break. Marking the document untrusted made no difference, and Jev lacks the reasoning-effort setting that raised attack costs in the comparison models.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+18
- Incentives72
- Confidence57
A dev.to team replaced its coding agent's LLM with mocks built to sabotage the repo. The first run let 21 adversarial commits through, and the fix was path and AST checks that run after the model call and before git commit.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+35
- Incentives60
- Confidence45
A LessWrong stress test reports that roughly 50K tokens of opposing finetuning beat 190M tokens of midtrained motivations, in an implementation its authors assembled from the best public description of the method.
Reality
- Evidence38
- Adoption22
- Hype gap+20
- Incentives55
- Confidence45
Strands Evals 1.0 added chaos testing and red teaming in June. One AWS builder pointed both at a fixture-fed copy of his cost-reporting agent, because Cost Explorer charges a cent per API request.
Reality
- Evidence40
- Adoption15
- Hype gap+30
- Incentives55
- Confidence35
Haize Labs sold red teaming and evaluation tooling to frontier labs and large enterprises. Its engineers are now an internal research group inside a Toronto holding company that buys small-business software.
Reality
- Evidence45
- Adoption30
- Hype gap+12
- Incentives70
- Confidence52
Editing refusal behaviour out of the weights means nothing in front of the model can put it back, and the harder question is what the hosted version buys when SaferAI found the unmodified predecessor already refused nothing.
Reality
- Evidence42
- Adoption20
- Hype gap+25
- Incentives78
- Confidence40
Researchers at Shanghai Jiao Tong University and Ant Group report up to 76.6% attack success against an agent memory system without ever touching the store, which moves the problem from filtering inputs to trusting stored state.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+32
- Incentives55
- Confidence42
The round doubles the company's lifetime funding and sells frontier labs their adversarial testing off the shelf. Bloomberg and the Israeli press named an investor the release did not.
Reality
- Evidence45
- Adoption58
- Hype gap+28
- Incentives76
- Confidence54
PlannerCritic's author tried to inject his own engine. The blocks arrived as feasibility verdicts rather than safety strings, which is a design worth copying and a limit worth reading closely.
Reality
- Evidence44
- Adoption12
- Hype gap+24
- Incentives72
- Confidence41
Fortinet says the price was immaterial to its business. For teams budgeting a standalone AI red-teaming and guardrails tool, that is the number that matters.
Reality
- Evidence34
- Adoption18
- Hype gap+32
- Incentives82
- Confidence58