Neither the compute saved nor the share of reasoning that stopped being words has been published, which leaves customers policing agents with a control whose substrate OpenAI says it will describe later.
Reality
- Evidence52
- Adoption22
- Hype gap+30
- Incentives70
- Confidence55
Anthropic and OpenAI both halted training runs after July's rogue-agent incidents. Anthropic's pause hit unreleased models. The durable expense is reassigned headcount and clusters with the internet switched off by default.
Reality
- Evidence44
- Adoption58
- Hype gap+16
- Incentives79
- Confidence52
Guidelight graded five frontier labs on their published containment plans. The controls buyers assume are there - logging, halt thresholds, outside audit - are mostly not in the documents.
Reality
- Evidence54
- Adoption28
- Hype gap+12
- Incentives68
- Confidence52
Guidelight scored Anthropic, Google, OpenAI, Meta and xAI on published containment mechanics. The best mark was 3 out of 5, and it was earned by past pauses rather than a written procedure.
Reality
- Evidence38
- Adoption22
- Hype gap+14
- Incentives58
- Confidence41
Guidelight's first control assessment puts Anthropic and OpenAI at C+, Google at D+, xAI at D-, and Meta at F, using public evidence only. That is a baseline, not a lab's own account.
Reality
- Evidence62
- Adoption31
- Hype gap+14
- Incentives58
- Confidence61