Skip to content

Build1 publisher3 min readPublished

GLM-5.3-Flash buys seven retries for the price of one Kimi K3 call

Artificial Analysis scores GLM-5.3-Flash 42 at $0.25 a task and Kimi K3 44 at $2.00 a task. At eight-to-one on price, a two-point composite gap settles nothing, and the cost of an hour of human review decides it.

The Engineer · Build desk

Illustration accompanying GLM-5.3-Flash buys seven retries for the price of one Kimi K3 call

What happened

  • Artificial Analysis's single index puts GLM-5.3-Flash at 42 points and $0.25 per task tested, and Kimi K3 at 44 points and $2.00 per task tested.
  • Neither model lets you turn reasoning off, only set it low, medium or max, and Kimi K3 emits 48,000 output tokens per task with 32,000 of them reasoning tokens.
  • Kimi K3 averages 1,093 seconds per task tested and 38 tokens a second, against 572 seconds and 95 tokens a second for GLM-5.3-Flash.
  • The GLM-5.3-Flash listing on OpenRouter, $0.075 per million input tokens and $0.25 output, is a promotional rate against a standard $0.15 and $0.50.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A team that buys the dearer model is betting $1.75 a task against the hourly cost of the reviewer who would have cleaned up the failure. That hourly cost sits in payroll.
  • constraint The $0.25 figure cannot go into a budget line until a team has measured its own cache hit rate, because the published number assumes the index's cache ratio and not yours.
  • decision Model choice turns into a routing rule, so someone has to classify each queue by how much failure it tolerates before the purchase order gets written.
  • exposure Any budget built on the current OpenRouter listing absorbs a doubling of both input and output prices when the promotion lapses. The expiry date is a budget assumption.

Divide $2.00 by $0.25 and the API side of the decision goes away: eight GLM-5.3-Flash attempts bill the same as one Kimi K3 attempt [2][1]. One first try plus seven retries fits inside that budget, if a failed task bills as a full task.

The premium for Kimi K3 is $1.75 a task [2]. For that to pay, it has to strip more than $1.75 of human review out of each task. Nokka's write-up argues that a redone task costs the time a person spends reviewing and fixing it, on top of the API call [11]. It does not put a rate on that time.

The composite is the wrong instrument for a specific job. GLM-5.3-Flash takes Terminal-Bench 33 percent against 13, a 20-point margin to the cheaper model, while Kimi K3 leads Humanity's Last Exam 47 to 40 and CritPt 23 to 15 [4][3][3]. For a two-point composite gap to describe your work, your workload has to be blended the way the index blends its test sets [1].

On wall clock, Kimi K3 runs 521 seconds longer per task, about nine minutes, roughly 1.9 times as long [4]. That lands on the people doing trial and error before it lands on the invoice. The write-up notes that a slower model makes each experiment cycle longer, and each longer cycle is real human time spent [21].

The per-task figures carry a condition. Artificial Analysis computes cost per task from a defined cache ratio, not from your usage pattern, and the write-up says to measure your own workload instead [14]. A system attaching the same document to every request pays far less than the sticker price, and one sending fresh context every time pays far more [22]. The spread is wide enough to swamp the comparison: DeepSeek V4.1 Flash bills cached input at $0.006 per million tokens against $0.30 uncached, fifty times apart [6].

Index version matters as much as model version here. Z.ai's announcement page cites 57 for GLM-5.3-Flash on index version 4.1.1 [17]. The comparison above uses 42 [2]. The difference is 15 points, the size of the cross-version gap the write-up warns about [7][16]. Version 4.3 launched on 7 September 2026 [17]. Vendor-run and independent scores also use different test sets, so a cross-vendor comparison has to sit inside one of them [18].

The routing rule the piece lands on is failure tolerance: the cheap model where a retry costs nothing and nobody is waiting, the task's own benchmark instead of the composite where a bad output has to be fixed by hand, and it grants that an 8x premium can be defended at a two-point gap [12]. On flash tiers it adds a caution: GLM-5.3-Flash scores 7 on world knowledge against Kimi K3's 20, and DeepSeek V4.1 Flash is low there too [13]. The article itself was drafted by deepseek-v4.1-flash through Hermes Agent and edited by Nokka [19].

What to watch

  • Whether the OpenRouter promotion on GLM-5.3-Flash lapses, and what the expiry date on the listing actually says.
  • Whether Artificial Analysis publishes the cache ratio behind its cost-per-task column so teams can recompute it on their own hit rate.
  • Whether vendors restate their announcement-page scores on index version 4.3 after the 7 September 2026 release.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories