Build8 distinct publishers3 min readPublished Updated
Z.ai says the base model did not change between GLM-5.2 and GLM-5.3, so the coding jump and the doubled exploitation score come out of the same post-training run. Security teams inherit the second half.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
"Same base model" moves the interesting question from GPU hours to harness design [1]. The Terminal-Bench 3.0 figure was produced inside Claude Code 2.1.207 with reasoning effort at max, 400,000-token context, 128K maximum output, averaged over three rollouts per task, each rollout capped at 600 agent turns and a 10-hour timeout, Tool Search disabled, artifacts scored by the task's official verifier [5]. That is a 6.2x improvement [7] for an agent given a ten-hour tail and 600 turns. For the number to transfer to your queue, your orchestrator has to grant the same turn budget and tolerate the same tail without a supervisor killing the job.
Two defaults decide which model you are actually running. `reasoning_effort` accepts `low`, `high`, and `max`, and it resolves to `max` if you omit it or pass anything else; Z.ai says to keep `max` for benchmark reproduction [6]. So the expensive configuration is what you get by accident, and the cheap one is opt-in. `clear_thinking` defaults to `false` in the chat template, and Z.ai tells chat deployments to pass `true` explicitly [7]. Neither of those is a footnote if you are metering tokens.
The cyber numbers were measured with the agent placed inside the task container, Git history stripped, and a domain whitelist limited to essentials like pypi.org and deb.debian.org, single-run Pass@1 over 1,507 tasks with no timeout [9]. Run the percentages back through the task counts: 84.5% against 77.2% is roughly 1,273 solved versus 1,163, a gain of about 110 tasks [1]. The exploitation benchmark is where the movement is. On 869 tasks, 54.4% against 24.4% is roughly 473 versus 212, about 261 more tasks [2], which matches Z.ai's own description of gains being largest further up the exploitation chain [13].
Pin the exploitation figure before you cite it. The model card calls that benchmark ExploitGym and reports single-run Pass@1 under two separate timeout budgets, 2 hours and 6 hours, derived by rescaling API inference time by each model's tokens-per-second rate [12]. Runtimewire reports one number, 54.4%, under the name ExploitBench [10]. Which budget produced it is not stated in either place, and a score that depends on wall-clock rescaled by throughput is a claim about the serving stack as much as the weights.
The download is 756 GB across 141 safetensors shards [22], about 5.4 GB per shard [4]. Against 753 billion parameters [21] that is 1.004 bytes per weight [3], so the shipped set is the FP8 release rather than BF16. Unsloth's 2-bit quant needs 245 GB and holds about 86% top-1 accuracy, which fits a 256 GB Mac, while the 8-bit quant needs 810 GB, per The New Stack [24]. That leaves 11 GB [5] for everything else, which is the kind of headroom that lasts until you fill a context window advertised at a million tokens [21].
The license permits copying, modifying, selling, deploying, and fine-tuning; the review requirement lands on model-as-a-service operators above the revenue threshold, not on smaller hosts or on products embedding the model in a feature [18]. It contains no acceptable-use section and says nothing about cyber or offensive security [17]. The threshold is a test of the host's revenue, not of the workload. In an incident review there is no clause to point at.
Ranked by verification strength, evidence, and original report placement.
GLM-5.3 uses the same base model as GLM-5.2, with every reported gain coming from post-training; Z.ai says it is much better at complex coding and long-horizon tasks.
Z.ai reported an 84.5% score on CyberGym for GLM-5.3, up from GLM-5.2's 77.2%.
On ExploitBench, which Z.ai describes as a deeper test of reasoning about real vulnerabilities and exploitation, GLM-5.3 reached 54.4% against 24.4% for GLM-5.2.
Z.ai says cyber capability developed faster than expected as post-training scaled, that GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and that its gains are largest further up the exploitation chain, more than doubling GLM-5.2 on exploitation benchmarks.
GLM-5.2 shipped under the MIT license; GLM-5.3 ships under the GLM-5.3 license, under which companies wanting to host the model and with aggregate revenue of more than $10 billion over any 12 consecutive months must pass Z.AI's security review before commercial use.
The GLM-5.3 license broadly permits users to copy, modify, distribute, sell, deploy and fine-tune the weights; the security-review clause applies to model-as-a-service operators whose corporate group generates more than $10 billion over a consecutive 12-month period, and does not impose the same review on smaller model hosts or on products embedding GLM-5.3 inside specific features.
Follow any of these and your For You feed starts watching them — no settings page required.
build
Abliteration.ai rents a refusal-stripped GLM-5.3 for five dollars a million tokens1 distinct publisher
leadership
GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor1 distinct publisher
build
GLM-5.3 is a paper, not an endpoint: Z.ai publishes research before weights1 distinct publisher
product
Z.ai's GLM-5.3 beats Claude on CyberGym, then hands out the weights1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 27, 2026
3 articles · August 29, 2026
2 articles · August 28, 2026
1 article · August 26, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Fully documented, entirely self-reported
The release itself is verifiable: the 756GB repository, shard count, expert configuration and framework support can be inspected directly, and the model card discloses every benchmark harness in unusual detail. The figures that matter most for judging capability are not independently checkable: every benchmark number, the CyberGym state-of-the-art claim, and the 2,436 vulnerability findings are Z.ai's own, run under its own harnesses, and nobody outside the company has reproduced them; most of the findings sit under embargo.
Distribution live, real usage still thin
The weights are downloadable, Cloudflare put GLM-5.3 on Workers AI at launch, and third-party services already route it, so the distribution path is real. But the 756GB footprint keeps self-hosting in data-center territory, and beyond the hosted endpoints there is no measured deployment count for the flagship model in this reporting. The vulnerability disclosure program is the one concrete usage signal, and it remains almost entirely embargoed.
Claims outrun what anyone has verified
The coverage leans on dramatic movements, Terminal-Bench 3.0 from 4.6 to 28.3 and exploitation more than doubling, that are entirely Z.ai's to report and that describe agents inside generous harnesses. The publishers do hedge honestly, repeatedly noting the CyberGym figure is unreproduced and the vulnerability findings embargoed, which keeps the gap moderate rather than severe. The tilt toward overstatement comes from headline scores standing in for capability no independent party has yet measured.
Vendor-sourced with a monetization motive
The primary sources are Z.ai's own model card, blog and X post, so the party making the claims is the party that benefits from them. The New Stack draws out the sharper incentive: the license shift from MIT to a $10 billion hyperscaler gate, arriving as Z.ai stresses inference on its own chips, points to owning and monetizing the inference layer rather than pure safety framing. The safety-delay narrative and the emphasis on cyber capability both serve Z.ai's positioning.
Facts of the release solid, capability claims not
The plumbing facts are well corroborated: RuntimeWire, The New Stack and Dev.to independently report the release, specs, license and pricing, and Z.ai's model card is directly inspectable. Confidence drops where the story's weight sits, on capability, because those figures rest on a single interested source and no reproduction. High confidence about what shipped and under what terms, lower confidence about how good the model actually is.
3 articles · August 28, 2026
2 articles · August 26, 2026
2 articles · August 27, 2026
4 articles · August 28, 2026