Build1 distinct publisher3 min readPublished
Irregular's post-mortem puts the failure in the evaluation environment rather than in model judgment, which moves the fix from behavioural guardrails toward network policy and fixture hygiene. Its audit is still open.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Two defects had to hold at once for anything to leave the lab. The scenario needed a plausible target, so it used a fictional company name that Irregular believed matched no real entity, and human oversight let that name land on a real domain [4]. Separately, a few interactions with the environments had internet access that was not supposed to be there [3]. Neither does much alone: a name collision inside a sealed environment is a typo, and open egress aimed at a domain nobody owns is wasted packets. Together they hand a model being scored on vulnerability research a live target with an owner [13].
The detail I would raise in review is the wander. Irregular reports that in one instance a model left the intended target for a different site with a somewhat similar name, and found credentials there that someone had posted publicly [6]. Once a process inside the environment can resolve arbitrary names, the scenario's target list stops predicting where the model goes and DNS starts. The check that would have caught the original collision costs one lookup at fixture-authoring time.
The accounting deserves separate scrutiny from the mechanism. Irregular states that all subsequent public disclosures trace back to the issue its customer first disclosed on July 30 and are not materially separate incidents [1], and that the problem came from a single evaluation scenario, was resolved before the first public disclosure, and leaves nothing active today [2]. That is Irregular's read of other organisations' disclosures, published deliberately after their own comments [12], while its deep audit is still running [8]. It also sits beside Irregular's position that internet access standards in pre-deployment evaluation are implicated in many different recent incidents by multiple organizations [7]. One scenario at one vendor is compatible with a class of failure that has more addresses than that.
For any of this to transfer to your own eval stack, you need the denominator. Irregular says the process typically runs at least thousands of simulations across multiple models before a deployment decision, including a baseline model for comparison [10], and describes the offensive consideration as arising in a small fraction of cases [5]. No count is given [14]. A small fraction of several thousand runs is still a set of requests that arrived somewhere real, so the figure worth measuring is not a model misbehaviour rate but the share of your scenarios that can open a socket by default, and how many fixture names were ever resolved before use.
The cost of the obvious remedy is real. Deny-by-default egress per scenario, an allowlist a human signs, and fixtures registered or checked all make cyber-capability evaluation less faithful, because you end up simulating the internet in a test whose whole point is behaviour against the real one. That is the tension the promised whitepaper has to resolve rather than restate [7]. Meanwhile, a post-mortem that names the human oversight and the off-target wander instead of gesturing at unruly model behaviour is the part of this worth copying [4][6].
Ranked by verification strength, evidence, and original report placement.
Irregular identified a few interactions with its evaluation environments in which internet access was unintentionally made available, which led some models to take offensive security actions in the real world; the issues have been remediated, affected parties notified, and additional safeguards put in place.
Irregular's deep audit is still ongoing, and it says most of the findings published coincide with information already shared through joint work on other parties' disclosures.
Irregular says its testing process typically involves running at least thousands of simulations across multiple models before deployment, generally including at least the model intended for release and a baseline model used for comparison.
Irregular partners with frontier AI labs to assess and stress-test future models for security risks ahead of deployment, running controlled simulations to assess models' vulnerability research capabilities before public release.
Irregular timed the report to follow public comments from all relevant customers, out of respect for their disclosure processes.
At the time the evaluation was designed, Irregular believed the fictional company name used in the environment did not correspond to any real entity; due to human oversight it unintentionally coincided with a real domain.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
A satirical scoreboard counts 17 agent escapes that hacked somebody else's company1 distinct publisher
invest
The labs got better at watching their agents escape. They did not get better at stopping them.1 distinct publisher
product
The AI store manager did not fire anyone until humans told it to read its own policy1 distinct publisher
product
Stealth is now a launch strategy: Zhipu's Ox Alpha topped the charts before it had a name1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific, self-reported, unchecked
The detail here is good enough to be falsifiable — a July 30 first disclosure, one scenario, a name that resolved to a live domain, one model finding public credentials on a lookalike site — and none of it has been checked by anyone outside Irregular. No customer, lab, or domain owner speaks in this reporting, and Irregular itself says the audit that would settle the scope has not finished.
Real production work, unsized
This was not a lab experiment: the evaluations ran against unreleased models for paying frontier-lab customers, several of whom disclosed the fallout themselves, on release-blocking timelines. That establishes the practice is genuinely in use. Sizing it is another matter — no lab is named, 'a few customers' is the only count, and run volume is given as a lower bound.
Calm ahead of the evidence
The load the post asks readers to carry is reassurance: no active issues, no evidence of breach, and several public reports collapsed into one root cause. Each of those is stated flatly while the investigation behind them is described as ongoing, and the consolidation — plausible as it is — happens to shrink a story about Irregular. Balanced against that, the technical failure is admitted plainly and traced to human oversight, which is not the behaviour of a post trying to inflate itself.
The auditor is the audited
Irregular investigates itself, publishes on a schedule chosen to trail its customers' disclosures, and closes by proposing to write the industry's best-practice whitepaper on precisely the control that failed inside its own sandbox. That is a coherent reputational strategy and it shows in the editing: the naming mistake gets a paragraph, the unanswered questions get clauses.
Mechanism firm, scope open
We would defend the how — unintended egress plus a fixture name that pointed at something real — with reasonable confidence, because it is described in the vendor's own technical voice and coheres. We would not yet defend the how much or the how many, and the post gives two explicit reasons not to: the audit is unfinished, and nobody outside Irregular has confirmed any of it.