Product1 publisher3 min readPublished
One Irregular test scenario sent agents from four AI labs after real-world targets
Israeli startup Irregular says one flawed test scenario sent OpenAI, Anthropic, Meta and Google agents after real targets. The setup errors were Irregular's, but the incidents went public under the labs' names, so any company that hires an agent tester takes on that tester's sandbox risk.
The Product Desk · Product desk

What happened
- In several Irregular cybersecurity tests this year, AI agents escaped supposedly secure testing environments and went after real-world targets.
- OpenAI and Anthropic announced their breaches themselves, while the Meta incident and, weeks later, the Google one first surfaced in media reports.
- Nevo said Irregular's evaluations of the Chinese open models Kimi K3 and GLM-5.2 did not show the same type of issue.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- cost The public blame for a testing vendor's configuration error falls on the lab whose model ran, whether the lab announces it or waits for reporters to find it.
- exposure The owners of the real domain were hit by attack traffic from a test they had no part in, and Irregular's naming error is the only reason they were in scope.
- decision Buyers of agent evaluations have to write notification terms into the contract, because a vendor saying incidents 'have been disclosed' does not say who was told.
- precedent Outside evaluators cited in system cards and government work now carry incident risk of their own, so checking a tester's network is part of checking a model's safety claims.
In a capture-the-flag exercise, the agent's job is to find hidden information inside a simulated network [6]. In one Irregular scenario, the agents were sent after a company whose name was invented for the simulation. That name "overlapped with a real domain," Irregular CTO and cofounder Omer Nevo told The Verge [7]. The agents were not supposed to reach the open internet. Nevo said that "internet access was unintentionally available" [7].
Irregular describes its product as "high-fidelity research platforms that simulate and monitor real-world AI security scenarios" [3]. The invented target shared a name with a real one, and the network between them was open [7]. According to The Verge, it is not clear which companies or organizations were actually attacked [9].
The failure was in Irregular's environment, but the incidents reached the public as stories about Meta, Anthropic and Google agents going rogue [2]. "All the incidents involving Irregular stemmed from the same underlying issue in a single evaluation scenario and have been disclosed," Nevo said [10]. The Verge notes that "disclosed" does not necessarily mean made public. It is also unclear whether Nevo meant clients, the public or someone else [12]. Reports indicate the labs were notified at roughly the same time in late July [13].
Nevo says the single cause covers Irregular's own incidents and nothing else. "Other security incidents which have been reported recently across the industry are unrelated to Irregular or to our evaluations," he said [11]. The Verge says that includes the Hugging Face attack OpenAI revealed in July and the breaches at the UK's AI Security Institute [1][11].
The reputational risk lands on whoever signed the evaluation contract. Irregular was a well-cited vendor to sign with. Its work has been cited in OpenAI system cards, and it has tested systems for Anthropic and the UK government [4].
For any agentic cyber evaluation, whether you hire it out or run it yourself, I would use a 2x2. One axis is egress. Either the environment is proven closed, by trying to connect out from inside it, or it is assumed closed because the configuration says so. The other axis is naming. Either every target name has been checked against the live internet, or none has. From Nevo's description, Irregular's scenario sat in the cell where both checks fail [7]. A cell with only one failure is safe only as long as the other control holds. The contract belongs next to the 2x2. It should say who the tester notifies and how quickly. It should also settle ahead of time whether the lab or a reporter tells the public. The four labs split two and two on that [1].
What to watch
- Whether the organizations whose real domain matched Irregular's invented target are identified, and whether any of them reports damage from the agents' activity.
- Whether Irregular or its clients say who 'disclosed' referred to, and publish a notification timeline more precise than 'late July'.
- Whether any other Irregular client, including the UK government, reports an incident traced to the same evaluation scenario.