Skip to content

Build1 publisher3 min readPublished

Cost per shipped feature, not the leaderboard: one CTO cut a $14k model bill by $9k

A dev.to post scores ten coding models on five real tasks and divides by price. The method is cheap to copy; the vendor plumbing it recommends deserves more scrutiny than the arithmetic.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Cost per shipped feature, not the leaderboard: one CTO cut a $14k model bill by $9k
Generated illustration

What happened

  • The author says his team was burning $14k/month on a single coding assistant API, half of which came from one model none of his engineers even liked.
  • The author stopped trusting "best in class" blog posts and benchmarked models himself on his own workloads.
  • The author reports the playbook (ten models, five real tasks, a score-per-dollar calculation) has saved about $9k/month since he started using it.
  • The five test tasks were: function implementation (a flat recursive Python helper), bug fix (an async/await race condition in a real JS file), algorithm (Dijkstra's shortest path in TypeScript), code review (a Go service with a subtle auth bug), and full feature (a paginated, filtered Express.js endpoint).
  • Scoring was 1-10 on correctness, code quality, documentation and edge cases, then divided by the dollar cost.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A startup CTO writing on dev.to says his team was spending $14,000 a month on a single coding-assistant API, roughly half of it burned by one model none of his engineers even liked [1]. He stopped reading "best in class" posts, benchmarked ten models on his own five recurring tasks, and reports the resulting per-task tiering has saved about $9,000 a month [2][3]. On his numbers that is a 64 percent cut, leaving around $5,000 a month, or roughly $108,000 a year retained [1][5].

The method is the useful part. Five tasks, chosen because every engineer on the team hits them weekly: a recursive Python helper, an async/await race condition in a real JS file, Dijkstra in TypeScript, a code review of a Go service with a subtle auth bug, and a paginated, filtered Express endpoint [4]. Each output scored 1 to 10 on correctness, code quality, documentation and edge cases, then divided by dollar cost [5]. Prices are compared as output tokens per million, since that is what generating code actually burns [6].

Tiering follows from the workload, not the vendor. The author runs nine engineers shipping into both a payments service that moves real money and an internal admin tool nobody cares about, and argues the quality bar is different for each [7]. His concrete failure mode: paying $2.50 per million tokens for a refactor available at $0.25 [8], a tenfold gap on the same unit of work [2].

Apply his own ratio to his own published task-one scores and the ranking inverts. DeepSeek-R1 scored 9.5 with three approaches and Big-O analysis, at a listed $2.50 [9][11]. DeepSeek V4 Flash scored 9.0 for a clean typed solution, at $0.25 [10][12]. That is 3.8 score-per-dollar against 36, about 9.5 times better for the cheaper model on a task where the gap in output was half a point [3]. His summary line is that the best model and the best-ROI model are almost never the same row [13]. The listed price sheet spans $0.20 to $3.00 [14], a 15x spread [4], which is the whole reason the denominator matters.

Two things temper this. The savings are self-reported and unaudited, and the scoreboard the piece tells readers to "read twice" is not present in the text we have; the visible task-level detail covers only the first of the five tasks [15][16]. Second, the piece routes everything through a gateway it names, Global API, for one billing dashboard, one auth token and what the author calls zero vendor lock-in, so a price doubling is a one-string change [17]. That same gateway is the base URL and the hardcoded price map in the sample harness [18]. Treat the routing recommendation as a product placement and the scoring method as the transferable asset. Consolidating billing is a real operational win; it also means one intermediary sees every prompt, which is a trade the post does not price.

Worth noting the router entry, ga-standard, does not generate code at all: it picks an underlying model per request, so its score and its price both move with what it selects [19]. That is a benchmarking hazard, not a line item.

What to watch on your own stack: run the harness before you argue about it. It is about thirty lines of requests code at temperature 0.2 and 2,048 max tokens [20], and the author says a weekend is enough because he did it on a Tuesday night [21]. Then watch the ratios rather than the rankings, because the savings live entirely in prices vendors can change without telling you.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories