Leadership1 distinct publisher3 min readPublished
A week-long GLM-5.3-Flash preview ran on tens of thousands of domestic chips at per-token cost Z.ai says matches mainstream Nvidia parts, but with no vendor named and no power or throughput figures, the result speaks to inference sourcing and not to training.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The memory budget explains most of what Z.ai actually built. GLM-5.3-Flash holds 320 billion parameters and activates 18 billion of them per token [7], and at launch the company measured three times less attention computation and a key-value cache 4.4 times smaller than GLM-5.3's [8]. Weights and the cache were then compressed further so a memory-limited part could hold more of the workload [6], and the serving engine was written specifically for chips whose binding limit is memory capacity and bandwidth [4]. On this evidence, per-token parity came substantially from shaping the workload to fit the silicon on hand, which is a narrower and more useful claim than domestic silicon matching Nvidia.
The distinction matters to a sourcing decision because the numbers that would settle it are the ones Z.ai withheld [3]. Utilization and throughput per accelerator would tell a buyer how many chips a given token rate consumes; power consumption would tell them the operating bill. Absent both, cost comparable to mainstream Nvidia GPUs [1] is a statement about one company's internal accounting. Z.ai also credits a GLM-5.3 infrastructure agent with helping write kernels and diagnose bottlenecks, "creating a feedback loop in which the model helped optimize the system serving the model itself" [18], which is a real engineering result and not a normalized comparison.
The prices, unlike the costs, can be checked. On Aug. 26, 2026 Z.ai listed Flash at $0.15 per million input tokens, $0.03 cached and $0.50 output, with a launch discount halving those rates through Sept. 9 [13]; GLM-5.3 was priced the same day at $1.40 input and $4.40 output [14]. That is 9.3 times cheaper on input and 8.8 times cheaper on output [1], or $0.075 and $0.25 on the discounted tier [2]. Z.ai reported an Artificial Analysis Intelligence Index 4.1.1 score of 57 at $0.045 per task, saying the same score previously cost about ten times more, which implies roughly $0.45 [15][4]. The score had not yet appeared on the Artificial Analysis leaderboard [15].
A skeptic's version is short: an unnamed chip, an unaudited fleet, and a cost claim with no denominator is marketing. That is fair on the cost figure. What the disclosure does establish is that a fleet of tens of thousands of domestic accelerators stayed up through a public preview under a custom stack [2], which is a scheduling and kernel achievement whoever made the chips. It buys something for anyone sourcing inference capacity inside China. It buys nothing for a training roadmap, and Z.ai does not claim the model was trained on those parts [17].
Supply is where the ceiling sits. SemiAnalysis puts CXMT's 2026 HBM output at enough for only 250,000 to 300,000 Ascend 910C-equivalent packages [11], against the 812,000 AI accelerators Huawei shipped in 2025 for a 20.3% share of a roughly four-million-unit market [12]. That ceiling covers 31% to 37% of last year's Huawei volume alone [3]. TrendForce projects domestic parts taking nearly 90% of China's high-end segment in 2026, up from 45% [16], but that is a share of a narrower category than the unit market Huawei's 20.3% is measured against, so the two figures are not in direct conflict.
DeepSeek settled this question the hard way: its R2 training run never completed reliably on Ascend hardware even with Huawei engineers on site, and the team went back to Nvidia for training while keeping Ascend for inference [10]. That split is the arrangement to plan against this quarter. Whether it survives the decade is a question about HBM output rather than about serving software.
Ranked by verification strength, evidence, and original report placement.
Against GLM-5.3, Z.ai measured three times less attention computation and a 4.4-fold smaller key-value cache at launch.
TrendForce projects domestic accelerators will take nearly 90% of China's high-end market in 2026, up from 45% in 2025.
Z.ai named no chip model or vendor and published no power consumption, exact throughput, utilization rate or normalized Nvidia comparison; the cost-parity assertion cannot be checked from the material released, and none of the serving results has been independently audited.
The technical disclosure describes a custom inference engine built on SGLang for chips constrained mainly by memory capacity and bandwidth.
Z.ai split multimodal encoding, prompt prefill and token-by-token decoding into separately scheduled worker pools, an Encode-Prefill-Decode design that lets operators add capacity where a workload is backing up instead of asking every accelerator to handle the full sequence.
Other measures compressed model weights and the key-value cache so a memory-limited chip could hold more of the workload.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed vendor disclosure, unverifiable core claim
Architecture, pricing and benchmark figures are specific and internally consistent, but the headline cost-parity result rests on a single vendor statement with no chip vendor named and no power, throughput, utilization or normalized Nvidia comparison, and no independent audit. Only one publisher covers the story, and the strongest checkable facts (pricing, parameter counts, market-share and HBM estimates) are peripheral to the claim being made.
Shipped and priced, deployment self-reported
The model is released, MIT-licensed and commercially priced with a live discount window, which are verifiable adoption facts. The scale claim — tens of thousands of domestic accelerators over a week-long preview — is disclosed only by the vendor, and no third-party operator, customer or leaderboard usage is documented.
Headline outruns disclosed measurement
The parity-with-Nvidia framing implies a validated hardware-substitution result, while the released material supports only that a memory-aware serving design plus a sparsely activated model ran a preview at low published prices. Scope is narrower than the framing (inference, not training), the benchmark figure is not yet on the leaderboard, and the supply ceiling on HBM constrains any generalization from one deployment.
Strong vendor and national-champion incentives
Z.ai is the sole source for the parity claim and benefits commercially from an aggressively low price sheet and politically from demonstrating viable domestic-chip inference, while withholding the vendor name and the metrics that would allow checking. Third-party supply and share numbers come from commercial research firms (SemiAnalysis, TrendForce) whose projections shape the same procurement debate, and the benchmark figure is vendor-selected rather than published by the benchmark operator.
Moderate — one well-caveated source
The lone publisher documents its claims specifically and flags its own verification gaps, which supports moderate confidence in what was said and priced. Confidence in the underlying hardware parity assertion is low: no second publisher, no audit, no named vendor, and the deployment scale cannot be independently observed.
build
SMIC's first $3 billion quarter comes with a wafer price increase attached1 distinct publisher
build
China's accelerator swap makes Cambricon supply, not export policy, your ship-date risk1 distinct publisher
build
H200s reach China at 2.5% of the order book, and Hong Kong holds the rest2 distinct publishers
build
The chokepoint moved: ABF film, not lithography, now caps China's accelerator output1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026