Build1 distinct publisher3 min readUpdated
Guidelight scored Anthropic, Google, OpenAI, Meta and xAI on published containment mechanics. The best mark was 3 out of 5, and it was earned by past pauses rather than a written procedure.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Containment is a runbook problem, and a missing runbook is easy to measure. The Guidelight rubric names the fields plainly: which of a model's accesses are revoked first, on what conditions it is allowed to keep working, and the point at which the system is switched off entirely [2]. Around those sit the housekeeping that makes any of it checkable, internal logs of model actions, monitoring for suspicious activity, a response to repeated violations, and independent verification of the control mechanisms [3]. None of that is research work. It is the same material a bank writes for a payments outage.
The answers the labs did give read like the first line of a plan. Anthropic's, according to Guidelight's review as reported by TechCrunch and relayed by mezha.net, is that if a model tries to evade oversight the company will assess the risk to determine whether containment is needed [8][14]. That commits to a judgement, not to a sequence. It does not say who holds authority to pull credentials, how quickly, or what the model retains while the assessment runs. Meta did not say whether an internal incident plan exists and pointed instead to its own risk assessment system [9], which is the pre-deployment genre standing in for the post-deployment one.
OpenAI's 3 out of 5 [4] is the more instructive score because of what earned it. Guidelight credited actual stoppages: internal deployments and training runs halted after safety incidents [5]. That is track record. The reviewers still found no formalised procedure in the public materials for a future case where control over a model is lost [6]. So the best-documented lab of the five has published its past reflexes rather than its future ones, and the ceiling for the whole group sits at 60 percent of a rubric any operations team would recognise [15], with the other four at or below that mark [16].
Steven Adler, Guidelight's chief scientist and a former OpenAI safety researcher, said he was surprised how little the companies said about how they would act in a very serious incident where a model went out of control, and that he hopes they have worked it through privately even where they have not said so publicly [11][12]. Private is the load-bearing word for anyone buying. Guidelight is clear that a low score reflects published material and is not evidence that internal safeguards are absent [10]. But an undisclosed trigger cannot be planned against, and Guidelight's own framing puts access limits, action monitoring and full shutdown at the centre of operational risk assessment as agentic systems spread through corporate environments [13]. The gap worth pressing in a procurement call is narrow and specific: which access comes off first, and who reads the log afterwards [2][3].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Guidelight AI Standards analysed the publicly available documents of Anthropic, Google, OpenAI, Meta and xAI and concluded that clear containment plans are almost absent from them.
Guidelight's assessment looked for plans describing which of a model's accesses must be revoked immediately, under what conditions it may be allowed to continue operating, and at what point the system must be fully shut down.
The assessment also covered internal logging of model actions, tracking of suspicious activity, response to repeated violations, and independent verification of control mechanisms.
OpenAI posted the best result of the five companies, 3 points out of 5.
Guidelight noted that OpenAI had already paused or terminated individual processes after safety-related incidents, including internal deployment and model training.
In OpenAI's public materials the researchers found no formalised procedure for a future case in which control over a model could be lost.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific scorecard, but one secondary source and an unpublished rubric
The factual core is concrete and attributable: a named assessor, five named labs, a stated rubric scope, one disclosed number (3 of 5), a ranking of lowest performers, on-record responses from Anthropic and Meta, and an explicit self-caveat. Against that, everything reaches the reader through a single Ukrainian-language item that credits TechCrunch, the rubric and per-company scores are not published or linked, there is no independent check of Guidelight's document reading, and Google, OpenAI and xAI are not quoted. That supports 'reported' rather than 'established'.
Published containment procedure is close to nonexistent across the five labs
Read as uptake of the practice the story is about — publishing an operational containment and shutdown procedure — the measured level is low by the assessor's own scoring: the ceiling across five frontier labs is 3 of 5, that leading mark rests on past ad hoc pauses rather than a written procedure, and the two lowest scorers either describe case-by-case risk assessment or decline to confirm an internal plan. Not zero, because OpenAI has demonstrably halted internal deployment and training runs after safety incidents. No enterprise-side adoption figures for agentic deployment are given, so this measures publication practice only.
Headline framing runs slightly ahead of a disclosure-only measurement
Mildly overstated. 'Almost no published plan for switching a model off' is a fair reading of the scores, but the assessed object is public documentation, and Guidelight explicitly says a low score does not mean internal safeguards are missing — a caveat the framing carries only in the final section. The gap is kept small because the story is a critical finding rather than a vendor promotion, the one disclosed number is stated plainly, and both named low scorers were given space to respond. The unpublished rubric and the single undisclosed-score ranking are what prevent a zero reading.
Standards body scoring the field it wants to standardise; subjects defending themselves
Legible and moderate. Guidelight AI Standards gains authority and relevance by publishing a scorecard that positions containment disclosure as a benchmark for operational risk assessment, and its chief scientist is a former OpenAI safety researcher now grading his ex-employer's field. The labs' quoted replies are self-interested in the opposite direction: Anthropic recasts the gap as risk-based judgement, Meta redirects to its existing risk framework. The publisher's incentive is aggregation — restating TechCrunch reporting for its own audience. None of this is concealed, which limits the score.
Direction credible, specifics unverifiable from this cluster
The overall direction — frontier labs publish little about isolating or shutting down a model that evades oversight — is coherent, internally consistent and reinforced by the labs' own non-committal responses. Confidence stays below midpoint because a single secondary item carries the whole cluster, only one of five scores is disclosed, the rubric is not available for inspection, three of five subjects are unquoted, and the assessor's caveat means the finding cannot speak to actual internal readiness.
product
OpenAI first, Anthropic and Meta last: a containment ranking of what labs admit in public1 distinct publisher
build
The best grade for controlling in-house AI agents is a C+, and buyers can now cite it2 distinct publishers
invest
The labs got better at watching their agents escape. They did not get better at stopping them.1 distinct publisher
build
First-turn evals test the safest part of your product, a 90,000-exchange audit finds1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
mezha.net
1 article · August 22, 2026