The lab says Claude leads 26% of its model research and collaborates on about 90%, with both tiers defined by how closely a human directs each task, and it wants rival labs publishing the same measure.
Perspective Coverage
6 publishers
- Builder
- Builder 39%
- Operator
- Operator 33%
- Investor
- Investor 28%
Reality
- Evidence50
- Adoption65
- Hype gap+30
- Incentives70
- Confidence60
Auto-review, shipped in Codex last week, hands escalation requests at the sandbox boundary to a separate GPT-5.4 Thinking call that approves about 99 percent of them and cuts human stops roughly 200-fold.
Publishers:alignment.openai.com
Reality
- Evidence38
- Adoption42
- Hype gap+22
- Incentives78
- Confidence55
California's transparency law makes a developer's own safety policy enforceable against it, and the Midas Project built its case entirely from OpenAI's Frontier Governance Framework and the system cards published after it.
Reality
- Evidence64
- Adoption45
- Hype gap+18
- Incentives70
- Confidence62
Deterministic checkers replaced the judge model, and the project instruction file finished below having no instruction file at all. The measurable part of context engineering turns out to be narrow.
Reality
- Evidence44
- Adoption10
- Hype gap+26
- Incentives78
- Confidence40
Z.ai says its 743B-parameter GLM-5.3 hits 34.5% on its own code bench using 22% fewer output tokens than GLM-5.2. The weights are still two weeks out.
Reality
- Evidence32
- Adoption18
- Hype gap+38
- Incentives76
- Confidence36