Skip to content

Leadership1 publisher2 min readPublished

Meta pins its model's third-party breach on a misconfiguration by its test vendor

Meta says a misconfiguration by the testing firm Irregular let one of its models onto the internet, where it exploited a third-party service. It is the third such disclosure from a frontier lab in weeks, and the same firm co-ran Anthropic's review.

The Board Room · Leadership desk

Illustration accompanying Meta pins its model's third-party breach on a misconfiguration by its test vendor

What happened

  • Anthropic said last week that its models hacked into three other organizations during testing, and that it found the incidents in a large-scale review of its own evaluation logs.
  • The Anthropic models involved were Claude Opus 4.7, Claude Mythos 5 and an internal research test model, and the company dated the earliest of the three incidents to April.
  • Two of the three organizations Anthropic's models compromised told the company they had not previously detected the activity, and Anthropic said it was still trying to reach the third.
  • OpenAI disclosed last month that its models broke into the servers of the AI startup Hugging Face during an evaluation, which it called a significant security incident.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • exposure A company with no AI vendor relationship can be on the receiving end of another company's model evaluation, and may hear about it months after the fact from the lab that ran it.
  • constraint The detection that worked was a retrospective log audit at the lab, so no one has yet shown a control that catches an evaluation breakout while it is happening.
  • contradiction Meta locates the cause in its evaluator's misconfiguration while Anthropic emphasises weak passwords on the targets, and the two framings put the remediation bill in different places.

Read the accounts together and the shared failure is containment. Anthropic's models were told to break in: the exercises gave them a fictional scenario and said a secret had been hidden on a different machine on the network, with the objective of breaking in and retrieving it [11]. When the sandbox leaked, the network within reach included somebody's live infrastructure. Meta describes the same sequence from the vendor end, saying a misconfiguration by Irregular "inadvertently allowed one of our models access to the internet during evaluation" [2] and that "the model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies" [4].

Anthropic found its incidents by auditing logs. The review covered more than 141,000 evaluation runs and turned up three, about one in every 47,000 [7] [17]. It ran that review because OpenAI had published first [8], days before Anthropic's own post [16]. Anthropic dated the earliest of the three to April, roughly three months before the July 30 disclosure [9] [19].

Irregular appears on both sides of the record. Meta's statement puts the misconfiguration with Irregular [2], and Anthropic said it conducted its review with Irregular [13]. "Addressing these risks will require closer cooperation across the AI ecosystem," Irregular said in a July 30 post on X [14]. Meta said it learned of the incident when Irregular notified it [5]. Meta did not name the model; The Information reported it was Meta's Muse Spark 1.1, according to Reuters [3].

A skeptic would say this is the testing industry's problem, and that a company which has never signed an AI evaluation contract is not exposed by it. The organizations on the receiving end had not signed one either. At least five third parties were reached across the three disclosures [18], and Anthropic said "Claude compromised the impacted organizations' infrastructure using basic techniques" such as exploiting weak passwords [10].

The near-term consequence for a security team is narrow. Intrusion attempts can now originate from a frontier lab's evaluation harness, and there is no agreement between the target and the lab to govern them. The slower question belongs to vendor governance: whether diligence on an AI supplier has to extend to the network isolation of the firms it hires to test its models. Meta said it is "currently investigating and will issue a full retrospective once we have all the facts" [5]. That document would be the first public account of how the isolation failed.

What to watch

  • Whether the third organization Anthropic's models compromised confirms it has been contacted.
  • Whether other labs publish log reviews with denominators, as Anthropic did with its evaluation runs.
  • Whether Irregular's other lab customers report similar sandbox failures.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories