Build1 distinct publisher3 min readPublished
Zhipu's open-weight MoE ties Claude Opus 4.8 on one composite index at roughly a fortieth of its list rate. At that price gap, the practical question stops being which model is better. It becomes how many attempts your router can afford.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The 1/40 figure is a per-token ratio measured against a list price, and the dev.to post carrying it says so before anyone else can [4][15]. Convert it into the unit that decides routing, which is cost per completed task. Apply the stated ratio to the post's own worked session and the same million-in, million-out job against Opus 4.8 list money runs about $60 [1]. A router can therefore attempt the cheap model forty times, or attempt it once and escalate thirty-nine failures, before it has spent what one flagship pass costs [2]. Retry tolerance is what the price gap actually buys.
Two things eat that allowance. The first is a discount you already hold: the post's caveat is explicit that the ratio is against list, so a negotiated Opus rate shrinks it before you write any code [15]. The second is the checker. Escalation only pays if something cheap decides the cheap answer was wrong. If the trigger is a failing test or a rejected schema, you keep most of the headroom. If the trigger is a person reading the output, review labour swamps both token bills and the ratio stops being the number that matters [16].
The price is possible because of the A18B in the name. Around 18B of 320B parameters are active per token, roughly 5.6 percent of the network doing arithmetic on any forward pass [1][5].
The post files this in the same commodity band as DeepSeek V4 Flash [10]. Measured against the rates in GLM's own worked example, that band contains a factor of four: input runs about 2.1 times DeepSeek's, output about 4.3 times [4]. Zhipu's claim that the overall bill still lands below an adjusted DeepSeek V4-Flash is the vendor's own [11], and the price table does not demonstrate it. For it to hold on your traffic, GLM would have to finish tasks in meaningfully fewer output tokens, because output is where the entire gap sits [3].
The tie with Opus is one composite score on one index, 57 against DeepSeek V4 Pro's 53 [2][3]. An aggregate tie is compatible with a large deficit in any single category, and the index does not exercise your tool schema, your prompt formatting, or your latency budget. For the number to transfer, your task mix has to resemble the index's eval mix, and your prompts have to survive a different tokenizer and a different tool-call dialect. The free tier is 200 requests a day [14], enough to sample your prompt set, not enough to load-test it.
One detail to read before planning around the promotion: in the post's arithmetic the discount window takes output from $1.20 to $0.60 per million and leaves input at $0.30 [5][6]. Input-heavy work gets nothing from it. Ten million input tokens is $3.00 either way [7][7], and repository re-reads and long carried agent state are exactly that shape.
The MIT licence promises an exit, though the hardware bill described below keeps that exit out of reach for most teams right now. At two bytes per parameter, 320B weights is 640GB of memory before any KV cache for a 1.04M-token context [1][12][6]. Eight 80GB cards hold the weights exactly and nothing else. The API price is the price, and the licence is insurance against that price changing.
In my context the tiering is worth building: send high-volume classification and document processing through an automatic validator, track the escalation rate, and let the flagship handle whatever fails [17]. I would not route single-shot work with no automatic check through the cheap tier, because there the 40x is a per-token figure and the real cost is the person who reads the answer. The budget under pressure is any budget that assumed all traffic goes through the flagship rate.
Ranked by verification strength, evidence, and original report placement.
GLM-5.3-Flash is Zhipu AI's MIT-open-source 320B-A18B mixture-of-experts model.
At standard international rates the post prices a session of 1M input and 1M output tokens as $0.3 input plus $1.2 output, or $1.50 per session.
In the limited-time half-price window the post prices the same 1M/1M session as $0.3 plus $0.6, or $0.90.
10M input tokens cost $3.00 at GLM-5.3-Flash's standard international rate.
Domestic GLM-5.3-Flash rates are 0.8 CNY per 1M input tokens and 2.8 CNY per 1M output tokens.
Domestic pricing is about 1/10th of GLM-5.3, the full-capability flagship tier, and about 1/20th during the limited-time half-price window.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test2 distinct publishers
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
product
Stealth is now a launch strategy: Zhipu's Ox Alpha topped the charts before it had a name1 distinct publisher
invest
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Checkable math on unchecked inputs
The arithmetic in this story is honest and reproducible; the inputs to it are not verifiable from anything published here. dev.to's post says "the sources" a dozen times without naming one, and the three numbers doing the most work — the index tie, Opus 4.8's list rate, the 62 trillion tokens — appear with no benchmark page, no price sheet and no platform data behind them. You can recompute $1.50 a session all day; you cannot confirm what it is a fortieth of.
One loud number, no receipts
Sixty-two trillion tokens in six days would be one of the fastest cold starts an inference platform has seen — and it reaches us as a clause in an opening sentence, unattributed to OpenRouter. Around it sit the things that are plainly real: a live price list in two currencies, a promotional window, a 200-request free endpoint. Product exists and is purchasable; the demand story is a stealth-launch anecdote nobody outside the post has confirmed, and no named team says it has moved traffic.
Headline outruns the ledger it rests on
This would score far worse if the post did not spend two paragraphs undercutting its own hook — list price versus negotiated rates, per-token versus per-task, run your own evals, budget against standard pricing. What keeps the gap positive is the framing that survives all that: one composite index score is treated as parity with a frontier flagship, thirty-nine retries are priced against a rate the reader never sees, and Zhipu's claim to a lower total bill than DeepSeek sits four inches from rates that run two to four times DeepSeek's per token, with no one pointing at the seam.
Launch-window framing, community byline
Follow who benefits from each number. The rate tables, the tenth-of-the-flagship comparison, the flagship-for-quality-extreme guidance and the beats-DeepSeek line all originate with the seller, and they land during a limited-time discount that rewards a memorable ratio. The counterweight: an individual developer's post owes Zhipu nothing, discloses no relationship either way, and volunteers the caveats a vendor would rather bury — which is also why the reader gets vendor framing without vendor accountability.
Trust the direction, not the decimal
Two facts cap this: one publisher, and every load-relevant figure sourced to an unnamed original. Cheap open-weight models compressing the price gap is a well-established direction, and the internal arithmetic here holds up, so the shape of the story is probably right. The specific quantities — a fortieth, 57 versus 53, sixty-two trillion — are exactly the kind that get rounded generously in a launch cycle, and nothing in our coverage would catch it if they had been.