Anthropic says Zhipu's freely downloadable GLM-5.3 built working V8 exploits in 50 of 410 tries, against 56 for its own restricted Claude Mythos Preview. With the weights public, its safeguards come off cheaply, so a lab that restricts its own model no longer keeps the capability out of reach.
Perspective Coverage
4 publishers
- Builder
- Builder 41%
- Operator
- Operator 38%
- Investor
- Investor 21%
Reality
- Evidence62
- Adoption30
- Hype gap+15
- Incentives72
- Confidence62
Anthropic says Zhipu AI's downloadable GLM-5.3 builds cyber exploits on its own, with safeguards that fail against simple attacks up to 100% of the time. It puts a capability once confined to gated frontier models within anyone's reach.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives75
- Confidence50
Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.
Perspective Coverage
5 publishers
- Builder
- Builder 58%
- Operator
- Operator 33%
- Investor
- Investor 9%
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives70
- Confidence55
The top rung of OpenAI's Preparedness Framework has now been reached by OpenAI, on a model it has not shipped, which moves AI-assisted exploitation out of argument and into a named vendor's published paperwork.
Perspective Coverage
5 publishers
- Builder
- Builder 28%
- Operator
- Operator 42%
- Investor
- Investor 30%
Reality
- Evidence35
- Adoption3
- Hype gap+30
- Incentives70
- Confidence55
The most capable cyber model OpenAI has built goes first to unnamed alpha testers guarding critical infrastructure, which means the buyers most eager to price it cannot bid, and what they get instead is an admission decision the company will not explain.
Publishers:fortune.com · uk.finance.yahoo.com Reality
- Evidence40
- Adoption5
- Hype gap+30
- Incentives70
- Confidence55
Astra scored a perfect 100% on OpenAI's own exploit-development benchmark, against 78.5% for GPT-5.6 Sol, and the shipped model's refusal to write proof-of-concept code is a policy the company has already said it will relax.
Perspective Coverage
4 publishers
- Builder
- Builder 30%
- Operator
- Operator 42%
- Investor
- Investor 28%
Reality
- Evidence35
- Adoption20
- Hype gap+45
- Incentives70
- Confidence55
Vendor benchmark tables are dated snapshots, and the GPT-6 Astra launch shows how much can move inside one afternoon without the headline score changing, which matters for anyone scoring a purchase off one.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence58
OpenAI's new model thinks repeatedly before it acts, and according to Manifold Security's CTO it usually does so without leaving the reasoning trace that agent audits read. Oversight moves to the buyer.
Perspective Coverage
18 publishers
- Builder
- Builder 32%
- Operator
- Operator 38%
- Investor
- Investor 30%
Reality
- Evidence52
- Adoption25
- Hype gap+45
- Incentives72
- Confidence60
Astra scored 100% on OpenAI's exploit-conversion benchmark with production safeguards switched off, and reached API and AWS customers the same week, with a refusal layer standing in for delay.
Perspective Coverage
8 publishers
- Builder
- Builder 36%
- Operator
- Operator 27%
- Investor
- Investor 37%
Reality
- Evidence45
- Adoption40
- Hype gap+40
- Incentives75
- Confidence60
Greg Brockman said AGI arrived with GPT-6 Astra on September 3. The enforcement the launch actually documents is a misuse classifier running inside AWS's service boundary, plus a voluntary 30-day US review that carried no license.
Perspective Coverage
18 publishers
- Builder
- Builder 40%
- Operator
- Operator 34%
- Investor
- Investor 26%
Reality
- Evidence60
- Adoption50
- Hype gap+45
- Incentives78
- Confidence66
OpenAI's spec lets gpt-6-astra take 922,000 input tokens, but requests above 272,000 move to higher long-context rates. Provisioning a model that drives a desktop, a shell and MCP servers starts with that threshold.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence40
DeAlignAI's downloadable FP8 build is the license working exactly as written, while its self-reported 320-of-320 HarmBench run remains unchecked by any outside researcher and measures compliance, not capability.
Reality
- Evidence46
- Adoption20
- Hype gap+18
- Incentives72
- Confidence55
ARC Prize ran GPT-6 Astra under two scaffolds and published both, 62.7% through its own minimal interface and 99.9% through OpenAI's Provider Adapter, which cost less. An eval that leaves the harness loose is scoring plumbing.
Reality
- Evidence68
- Adoption55
- Hype gap+64
- Incentives72
- Confidence62
Z.ai says the base model did not change between GLM-5.2 and GLM-5.3, so the coding jump and the doubled exploitation score come out of the same post-training run. Security teams inherit the second half.
Perspective Coverage
8 publishers
- Builder
- Builder 45%
- Operator
- Operator 31%
- Investor
- Investor 24%
Reality
- Evidence55
- Adoption40
- Hype gap+18
- Incentives75
- Confidence70
Reaching GPT-6 Astra costs a few days of patience rather than a $200 seat, so the gate worth budgeting against is Astra Pro, which Plus plans do not get, plus the credits sold on top of existing allowances.
Reality
- Evidence40
- Adoption25
- Hype gap+35
- Incentives70
- Confidence55
ARC Prize put GPT-6 Astra at 62.7% against OpenAI's 99.9%, and the third-party composite index has it 0.3 points above the model it replaces, which leaves the 20% safety compute overhead as the clearest number in the launch.
Reality
- Evidence57
- Adoption42
- Hype gap+58
- Incentives74
- Confidence54
The 100% ExploitBench score and two fresh V8 zero-days are OpenAI's own numbers, one of them still unverified, but the "critical" designation is a dated document that every agent deployer's controls will now be read against.
Reality
- Evidence28
- Adoption22
- Hype gap+58
- Incentives82
- Confidence34
OpenAI says Astra is the first model in any risk domain to reach Critical, on a score no outsider has seen. Buyers weighing today's releases have to grade the grader before they grade the model itself.
Reality
- Evidence46
- Adoption38
- Hype gap+30
- Incentives78
- Confidence50
Astra found the flaws while working through a twenty-vulnerability internal test, and OpenAI's answer is to ration offensive capability to vetted defenders rather than sell it by the token. That reprices exploit labour.
Reality
- Evidence33
- Adoption28
- Hype gap+34
- Incentives76
- Confidence40
OpenAI says Astra can find unknown flaws and chain them into working exploits without a human guiding each step. The evidence published so far is one saturated public benchmark plus a 20-vulnerability internal set, in which the model found two zero-days of its own.
Publishers:openai.com
Reality
- Evidence30
- Adoption10
- Hype gap+35
- Incentives78
- Confidence45