Build1 publisher3 min readPublished
GLM-5.3-Flash buys seven retries for the price of one Kimi K3 call
Artificial Analysis scores GLM-5.3-Flash 42 at $0.25 a task and Kimi K3 44 at $2.00 a task. At eight-to-one on price, a two-point composite gap settles nothing, and the cost of an hour of human review decides it.
The Engineer · Build desk

What happened
- Artificial Analysis's single index puts GLM-5.3-Flash at 42 points and $0.25 per task tested, and Kimi K3 at 44 points and $2.00 per task tested.
- Neither model lets you turn reasoning off, only set it low, medium or max, and Kimi K3 emits 48,000 output tokens per task with 32,000 of them reasoning tokens.
- Kimi K3 averages 1,093 seconds per task tested and 38 tokens a second, against 572 seconds and 95 tokens a second for GLM-5.3-Flash.
- The GLM-5.3-Flash listing on OpenRouter, $0.075 per million input tokens and $0.25 output, is a promotional rate against a standard $0.15 and $0.50.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A team that buys the dearer model is betting $1.75 a task against the hourly cost of the reviewer who would have cleaned up the failure. That hourly cost sits in payroll.
- constraint The $0.25 figure cannot go into a budget line until a team has measured its own cache hit rate, because the published number assumes the index's cache ratio and not yours.
- decision Model choice turns into a routing rule, so someone has to classify each queue by how much failure it tolerates before the purchase order gets written.
- exposure Any budget built on the current OpenRouter listing absorbs a doubling of both input and output prices when the promotion lapses. The expiry date is a budget assumption.
Divide $2.00 by $0.25 and the API side of the decision goes away: eight GLM-5.3-Flash attempts bill the same as one Kimi K3 attempt [2][1]. One first try plus seven retries fits inside that budget, if a failed task bills as a full task.
The premium for Kimi K3 is $1.75 a task [2]. For that to pay, it has to strip more than $1.75 of human review out of each task. Nokka's write-up argues that a redone task costs the time a person spends reviewing and fixing it, on top of the API call [11]. It does not put a rate on that time.
The composite is the wrong instrument for a specific job. GLM-5.3-Flash takes Terminal-Bench 33 percent against 13, a 20-point margin to the cheaper model, while Kimi K3 leads Humanity's Last Exam 47 to 40 and CritPt 23 to 15 [4][3][3]. For a two-point composite gap to describe your work, your workload has to be blended the way the index blends its test sets [1].
On wall clock, Kimi K3 runs 521 seconds longer per task, about nine minutes, roughly 1.9 times as long [4]. That lands on the people doing trial and error before it lands on the invoice. The write-up notes that a slower model makes each experiment cycle longer, and each longer cycle is real human time spent [21].
The per-task figures carry a condition. Artificial Analysis computes cost per task from a defined cache ratio, not from your usage pattern, and the write-up says to measure your own workload instead [14]. A system attaching the same document to every request pays far less than the sticker price, and one sending fresh context every time pays far more [22]. The spread is wide enough to swamp the comparison: DeepSeek V4.1 Flash bills cached input at $0.006 per million tokens against $0.30 uncached, fifty times apart [6].
Index version matters as much as model version here. Z.ai's announcement page cites 57 for GLM-5.3-Flash on index version 4.1.1 [17]. The comparison above uses 42 [2]. The difference is 15 points, the size of the cross-version gap the write-up warns about [7][16]. Version 4.3 launched on 7 September 2026 [17]. Vendor-run and independent scores also use different test sets, so a cross-vendor comparison has to sit inside one of them [18].
The routing rule the piece lands on is failure tolerance: the cheap model where a retry costs nothing and nobody is waiting, the task's own benchmark instead of the composite where a bad output has to be fixed by hand, and it grants that an 8x premium can be defended at a two-point gap [12]. On flash tiers it adds a caution: GLM-5.3-Flash scores 7 on world knowledge against Kimi K3's 20, and DeepSeek V4.1 Flash is low there too [13]. The article itself was drafted by deepseek-v4.1-flash through Hermes Agent and edited by Nokka [19].
What to watch
- Whether the OpenRouter promotion on GLM-5.3-Flash lapses, and what the expiry date on the listing actually says.
- Whether Artificial Analysis publishes the cache ratio behind its cost-per-task column so teams can recompute it on their own hit rate.
- Whether vendors restate their announcement-page scores on index version 4.3 after the 7 September 2026 release.