Build1 publisher3 min readPublished
DeepSeek 4.1 Flash undercuts a $20 coding seat only for teams running about one build a month
DeepSeek 4.1 Flash finished a metered Ship-Bench build for $15.04 in per-token fees and scored lowest at code review. Developers who run more than about one and a third full builds a month would still pay less on a $20 seat.
The Engineer · Build desk

What happened
- DeepSeek 4.1 Flash's best scores came where its tester expected weakness: Architect at 98.3, UX at 98.6 and Planner at 94.5.
- It passed all five phases and beat an earlier Grok 4.5 run, which averaged 92.0, in every role.
- Implementation alone cost $13.43 of the total, and the full run used roughly 412 million tokens at high effort.
- The finished app passed 492 unit tests at 95.46% line coverage plus 28 Playwright end-to-end journeys.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A developer who runs two full builds a month would spend $30.08 on tokens, so the $20 seat stays cheaper for anyone above about 1.33 builds.
- decision Moving review to a stronger reasoning model, as the author proposes, adds its price to one phase; every phase outside implementation cost $1.61 combined at Flash rates.
- contradiction The author says the Opus-class judge overrated an overreaching architecture spec at 98.3, so the 94.5 average partly reflects that judge's preference for exhaustive output.
- constraint The evidence is one build of a simplified knowledge base app, so the $20 comparison holds only for teams whose monthly use resembles that build.
"The per-token price is low, but the model makes it back in volume," the author of the Ship-Bench run wrote [4]. Implementation alone cost $13.43 of the $15.04 total [3]. That is about 89% of the bill [1]. Architecture, UX, planning and review together came to $1.61 [2].
Spread over roughly 412 million tokens [3], the run works out to about 3.7 cents per million tokens [3]. The post does not split that count into input, output and cached reads, so the blended rate cannot be checked against a price sheet. The run was also metered at high effort [3]. A lower effort setting or a different task would change both the token count and the bill.
A $20 seat's budget covers about 1.33 of these builds a month [4]. The author lands in the same place, more conservatively: pay-as-you-go beats the subscription "at roughly one full build per month or less" [5]. Two builds cost $30.08 [5]. For the comparison to transfer, a developer's month on the seat has to look like one end-to-end build of the simplified knowledge base app the series uses [9]. This run measured one such build, on the same machine and task as earlier entries [19].
The author expected a Flash model to be decent at coding and mediocre at architecture, design and planning [16]. "My hypothesis was backwards," the author wrote [16]. Architect scored 98.3, UX 98.6 and Planner 94.5, with Reviewer lowest at 85 [2], against a five-role average of 94.5 [1]. An Opus-class LLM judge assigned those scores [8], and the author disputes the Architect number. The spec pinned even minor packages to exact versions instead of semver ranges, listed the repo layout file by file, and pre-decided a styling and design system that belongs to the UX phase [14]. "I would score this lower than 98.3 given the overzealous nature of the output," the author wrote [13]. The Planner shipped 8 iterations against a gate of 3 to 5, and the judge recorded it as a pass with deviation [15].
The verification habit in that spec is good engineering. The 2,430-line document confirmed seven of its claims by running commands, including a reproducible ERESOLVE that justified pinning TypeScript 6 against the newer 7 [11]. The judge's only deduction was for a postinstall step that downloads roughly 400 MB of Playwright browsers [12]. I would have taken those points too.
Review is where the author would change the setup, with "a stronger reasoning model in the reviewer chair" [5]. An earlier Grok 4.5 run showed the same contour, strong early phases and a dip at Reviewer [17]. GitHub Copilot Desktop was the harness, with OpenRouter supplying the model [6]. The author says that harness is more permissive about tool approvals than the CLI variants used before, so the hands-off operator experience is partly a harness property [7]. Running the model locally instead needs roughly 512 GB of memory at 4-bit quantization [18].
What to watch
- A rerun of DeepSeek 4.1 Flash in a stricter CLI harness, which would separate model behaviour from Copilot Desktop's permissive tool approvals.
- A mixed run with a stronger reasoning model as Reviewer, with that phase's score and cost reported separately.
- A split of the 412 million tokens into input, output and cached reads, which would show whether the roughly 3.7 cents per million holds on other tasks.