Build1 distinct publisher2 min readPublished
Z.ai's new model activates 18B parameters per token and lists at a tenth of GLM-5.2's price. The card publishes its evaluation settings and, in the text supplied, not its scores.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The design choice worth reading here is the attention stack, not the headline parameter count. Z.ai says this is the first model in the GLM series to combine sparse and linear attention, and that the purpose of the combination is long-context serving cost while keeping precise long-context behaviour [5]. Around it sit Manifold-Constrained Hyper-Connections for scaling efficiency [6], a newly trained base model [4], and a 30T-token multimodal pre-training corpus [7]. On the compute side the arithmetic is plain: 18B of 320B works out at 5.6% of the weights active per token, with 94.4% untouched on any given forward pass [1]. That ratio is what a tenth-of-the-price API listing rests on [3][9], and it is a statement about work per token rather than about everything else on the invoice.
The evaluation footnotes are the most useful part of the card, because they describe the conditions under which the coding claim was measured. NL2Repo ran at 1M context with 64k new tokens [11]. DeepSWE ran in the mini-swe-agent harness at 400K context with a six-hour timeout [13]. Terminal-Bench 2.1 ran inside Claude Code 2.1.207 with 65,536-token generations and the same six-hour ceiling [14]. HLE with tools used a 300,000-token window plus a context management strategy, with GPT-5.6-luna at medium acting as judge [10].
Not all of those numbers are self-administered. GDPval-AA v2 was scored by Artificial Analysis [15], and Toolathlon Verified went through the official evaluation service, reported as pass@1 averaged over three runs [16]. The rest sit inside the vendor's own harness choices, which is ordinary practice, and also the reason the judge model and the context strategy matter as much as the resulting figure.
One footnote is a deployment note wearing benchmark clothes. NL2Repo judging combined rule-based and LLM-based checks specifically to catch unauthorized pip or curl operations [12]. A team writes that check because the model reaches for the network while it works, which tells you something about the sandbox you need before you run it against a real repository.
The scores themselves are not in the card text supplied here, only the settings that produced them [18]. That leaves three things in the pitch: fewer FLOPs per token, cheaper long-context serving, and parity with Claude Opus 4.8 on coding and agentic benchmarks [3]. The first is arithmetic, the second is measurable on rented hardware with SGLang or vLLM [8], and the third is currently a sentence.
Ranked by verification strength, evidence, and original report placement.
DeepSWE was run using the mini-swe-agent harness at temperature 0.95, top_p 1.0, a six-hour timeout and 400K context.
Terminal-Bench 2.1 was evaluated in Claude Code 2.1.207 at temperature 1.0, top_p 1, max_new_tokens 65536, with a six-hour timeout.
On GDPval-AA v2, models are evaluated by Artificial Analysis.
Toolathlon Verified results were obtained via the official evaluation service and reported as pass@1 averaged over three independent runs.
The supplied model card text gives evaluation configurations for HLE with tools, NL2Repo, DeepSWE, Terminal-Bench 2.1, Toolathlon Verified, AutomationBench, GDPval-AA v2 and BabyVision, but contains no benchmark scores.
GLM-5.3-Flash is introduced as the first natively multimodal model in the GLM-5 series.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed first-party methodology, zero published results
The single source is the vendor's own model card, which is primary and unusually specific about architecture and evaluation configuration (judge model, context limits, harness versions, benchmark patch level). But the supplied text contains no benchmark scores, no prices and no independent measurement, so the substantive performance and cost assertions cannot be checked from the cluster.
Availability documented, uptake unmeasured
Adoption evidence is limited to availability the vendor controls: four documented serving stacks and a hosted API on release day. The Hugging Face download counter is blank and no third-party deployment, customer or usage disclosure appears, so there is no observed uptake.
Headline outruns the published numbers
The card asserts it beats GLM-5.2 across benchmarks and real-world workloads at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic tasks, while the supplied text supplies no scores and no prices. Sparsity framing ('18B active') also understates the 320B-weight residency an operator must fund. The gap is moderate rather than extreme because the methodology disclosure is genuinely detailed and the card points to a blog post and technical report.
Vendor launch page selling its own API
The only source is authored by the model's developer, published on its own model card, framing a competitive win against its prior model and a frontier rival, and linking to the paid API platform where the model is sold. Benchmark selection, harness choice and judge model are all vendor-chosen, and no adversarial or independent voice is present in the cluster.
Primary document, single vendor voice
Confidence in what the card says is high because the source is the authoritative primary artifact quoted verbatim; confidence in the underlying performance and price picture is low because there is one publisher, no scores in the supplied text and no independent corroboration.
build
Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design2 distinct publishers
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly1 distinct publisher
build
Four frontier models in four days, and the cheapest number in your agent plan has an expiry date1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026