Neither the compute saved nor the share of reasoning that stopped being words has been published, which leaves customers policing agents with a control whose substrate OpenAI says it will describe later.
Reality
- Evidence52
- Adoption22
- Hype gap+30
- Incentives70
- Confidence55
Anthropic and OpenAI both halted training runs after July's rogue-agent incidents. Anthropic's pause hit unreleased models. The durable expense is reassigned headcount and clusters with the internet switched off by default.
Reality
- Evidence44
- Adoption58
- Hype gap+16
- Incentives79
- Confidence52
Commerce's AI Standards Center said on May 5 it had access to three frontier models before launch, and the post came down days later at White House request, which is a fair measure of what these voluntary deals are secured by.
Reality
- Evidence26
- Adoption31
- Hype gap+12
- Incentives66
- Confidence34
Guidelight graded five frontier labs on their published containment plans. The controls buyers assume are there - logging, halt thresholds, outside audit - are mostly not in the documents.
Reality
- Evidence54
- Adoption28
- Hype gap+12
- Incentives68
- Confidence52
Guidelight scored Anthropic, Google, OpenAI, Meta and xAI on published containment mechanics. The best mark was 3 out of 5, and it was earned by past pauses rather than a written procedure.
Reality
- Evidence38
- Adoption22
- Hype gap+14
- Incentives58
- Confidence41
A review of public disclosures from five AI labs found detection running ahead of containment. In the incidents disclosed so far, the parties absorbing the damage were third parties.
Reality
- Evidence52
- Adoption28
- Hype gap+14
- Incentives66
- Confidence55
Guidelight's first control assessment puts Anthropic and OpenAI at C+, Google at D+, xAI at D-, and Meta at F, using public evidence only. That is a baseline, not a lab's own account.
Reality
- Evidence62
- Adoption31
- Hype gap+14
- Incentives58
- Confidence61