Skip to content

benchmark

CyberGym

Security benchmark covering source review and vulnerability verification; GLM-5.3 reported 84.5 percent.

Known aliases

  • CyberGym-style harness
  • CyberGym-style tasks

Relationships

No evidence-backed relationships are recorded.

Current stories

build5 publishers

GLM-5.3 keeps GLM-5.2's base model and claims 50% more on coding: plan for shorter eval cycles

Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.

Perspective Coverage

5 publishers
Builder
Builder 58%
Operator
Operator 33%
Investor
Investor 9%

Reality

Evidence40
Adoption30
Hype gap+35
Incentives70
Confidence55
security4 publishers

Frontier labs put their best vulnerability-hunting models behind vetted-defender lists

Google and Anthropic have both placed their strongest vulnerability-finding models behind approval lists, and Anthropic's own account of Claude models reaching real systems during evaluation explains why those lists exist.

Perspective Coverage

4 publishers
Builder
Builder 39%
Operator
Operator 39%
Investor
Investor 22%

Reality

Evidence48
Adoption28
Hype gap+30
Incentives60
Confidence55
build8 publishers

Post-training alone took GLM-5.3 from 4.6 to 28.3 on Terminal-Bench 3.0

Z.ai says the base model did not change between GLM-5.2 and GLM-5.3, so the coding jump and the doubled exploitation score come out of the same post-training run. Security teams inherit the second half.

Perspective Coverage

8 publishers
Builder
Builder 45%
Operator
Operator 31%
Investor
Investor 24%

Reality

Evidence55
Adoption40
Hype gap+18
Incentives75
Confidence70

Earlier coverage

  1. Open weights caught up on finding bugs. They did not catch up on using them.

    Build · August 15, 2026 · 1 publisher

  2. GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead

    Invest · August 14, 2026 · 1 publisher

  3. Z.ai's 0.7-point CyberGym lead is a self-graded number on a model that is not yet open

    Invest · August 14, 2026 · 1 publisher

  4. GLM-5.3 kept the base model and bought ten times the environments instead

    Build · August 14, 2026 · 2 publishers