Skip to content

Build1 publisher3 min readPublished

DeepSeek 4.1 Flash undercuts a $20 coding seat only for teams running about one build a month

DeepSeek 4.1 Flash finished a metered Ship-Bench build for $15.04 in per-token fees and scored lowest at code review. Developers who run more than about one and a third full builds a month would still pay less on a $20 seat.

The Engineer · Build desk

Illustration accompanying DeepSeek 4.1 Flash undercuts a $20 coding seat only for teams running about one build a month

What happened

  • DeepSeek 4.1 Flash's best scores came where its tester expected weakness: Architect at 98.3, UX at 98.6 and Planner at 94.5.
  • It passed all five phases and beat an earlier Grok 4.5 run, which averaged 92.0, in every role.
  • Implementation alone cost $13.43 of the total, and the full run used roughly 412 million tokens at high effort.
  • The finished app passed 492 unit tests at 95.46% line coverage plus 28 Playwright end-to-end journeys.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A developer who runs two full builds a month would spend $30.08 on tokens, so the $20 seat stays cheaper for anyone above about 1.33 builds.
  • decision Moving review to a stronger reasoning model, as the author proposes, adds its price to one phase; every phase outside implementation cost $1.61 combined at Flash rates.
  • contradiction The author says the Opus-class judge overrated an overreaching architecture spec at 98.3, so the 94.5 average partly reflects that judge's preference for exhaustive output.
  • constraint The evidence is one build of a simplified knowledge base app, so the $20 comparison holds only for teams whose monthly use resembles that build.

"The per-token price is low, but the model makes it back in volume," the author of the Ship-Bench run wrote [4]. Implementation alone cost $13.43 of the $15.04 total [3]. That is about 89% of the bill [1]. Architecture, UX, planning and review together came to $1.61 [2].

Spread over roughly 412 million tokens [3], the run works out to about 3.7 cents per million tokens [3]. The post does not split that count into input, output and cached reads, so the blended rate cannot be checked against a price sheet. The run was also metered at high effort [3]. A lower effort setting or a different task would change both the token count and the bill.

A $20 seat's budget covers about 1.33 of these builds a month [4]. The author lands in the same place, more conservatively: pay-as-you-go beats the subscription "at roughly one full build per month or less" [5]. Two builds cost $30.08 [5]. For the comparison to transfer, a developer's month on the seat has to look like one end-to-end build of the simplified knowledge base app the series uses [9]. This run measured one such build, on the same machine and task as earlier entries [19].

The author expected a Flash model to be decent at coding and mediocre at architecture, design and planning [16]. "My hypothesis was backwards," the author wrote [16]. Architect scored 98.3, UX 98.6 and Planner 94.5, with Reviewer lowest at 85 [2], against a five-role average of 94.5 [1]. An Opus-class LLM judge assigned those scores [8], and the author disputes the Architect number. The spec pinned even minor packages to exact versions instead of semver ranges, listed the repo layout file by file, and pre-decided a styling and design system that belongs to the UX phase [14]. "I would score this lower than 98.3 given the overzealous nature of the output," the author wrote [13]. The Planner shipped 8 iterations against a gate of 3 to 5, and the judge recorded it as a pass with deviation [15].

The verification habit in that spec is good engineering. The 2,430-line document confirmed seven of its claims by running commands, including a reproducible ERESOLVE that justified pinning TypeScript 6 against the newer 7 [11]. The judge's only deduction was for a postinstall step that downloads roughly 400 MB of Playwright browsers [12]. I would have taken those points too.

Review is where the author would change the setup, with "a stronger reasoning model in the reviewer chair" [5]. An earlier Grok 4.5 run showed the same contour, strong early phases and a dip at Reviewer [17]. GitHub Copilot Desktop was the harness, with OpenRouter supplying the model [6]. The author says that harness is more permissive about tool approvals than the CLI variants used before, so the hands-off operator experience is partly a harness property [7]. Running the model locally instead needs roughly 512 GB of memory at 4-bit quantization [18].

What to watch

  • A rerun of DeepSeek 4.1 Flash in a stricter CLI harness, which would separate model behaviour from Copilot Desktop's permissive tool approvals.
  • A mixed run with a stronger reasoning model as Reviewer, with that phase's score and cost reported separately.
  • A split of the 412 million tokens into input, output and cached reads, which would show whether the roughly 3.7 cents per million holds on other tasks.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories