Skip to content

Build1 publisher3 min readPublished Updated

Your token ratio, not the leaderboard, decides which model is cheap

A dev.to price walkthrough shows two models swapping places by 17% and 42% on the same list prices. For that pair, the crossover sits at ten input tokens per output token.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Your token ratio, not the leaderboard, decides which model is cheap
Generated illustration

What happened

  • A post on dev.to titled "The cheap model is only cheap for half your tasks" argues that there is no such thing as a cheap model, only a model that is cheap for the shape of your traffic, and that the ranking reorders when the shape changes.
  • A token is roughly three quarters of a word; models bill per million tokens. Input tokens are what you send (prompt, files, chat history); output tokens are what the model writes back.
  • Input and output tokens have different prices, the gap between them is not the same for every model, and every vendor publishes the ratio between the two numbers without commenting on it.
  • The prices quoted in the post are list prices per 1M tokens, snapshot taken 12 Aug 2026.
  • claude-haiku-4-5 list price: $1.00 per 1M input tokens, $5.00 per 1M output tokens.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post on dev.to argues that there is no cheap model, only a model that is cheap for the shape of your traffic, and that the ranking reorders when the shape changes [1]. That is worth ten minutes of your time because every model has two prices, input and output, and the ratio between them differs by vendor without anyone drawing attention to it [4].

The arithmetic is small enough to check. Input tokens are what you send, output tokens are what comes back, and both are billed per million [3]. At the list prices the author snapshotted on 12 Aug 2026, claude-haiku-4-5 was $1.00 in and $5.00 out; grok-4.3 was $1.25 in and $2.50 out [5][6][7]. Haiku therefore charges five times as much to write as to read, Grok twice [1]. The author notes that a 2x multiple is unusual where most vendors sit at five or six [8].

Run the two through a classification job of 4,000 input tokens and 50 output tokens per request, 1,000 requests: Haiku costs $4.25, Grok $5.12 [9]. Haiku wins by 17% [10]. Now code generation, 1,500 in and 2,500 out: Haiku costs $14.00, Grok $8.12, and Grok wins by 42% [11][12]. Note that the 42% is measured against the more expensive option; expressed as a premium over the cheaper one, Haiku costs 72% more on that job [3].

Set the two cost equations equal and the crossover for this pair falls at exactly 10 input tokens per output token [2]. Above 10:1, Haiku. Below, Grok. That single number does more work than any leaderboard, and it explains why the source's own RAG shape, 8,000 in and 700 out, is nearly a coin toss: at 11.4:1 Haiku costs $11.50 and Grok $11.75, a 2% gap that should be decided on output quality instead [4].

The reordering also runs upward through the tiers. gemini-3.5-flash at $1.50/$9.00 costs $18.30 on that RAG workload; claude-sonnet-5 at $2.00/$10.00 costs $23.00 [13][14]. Sonnet is 26% more expensive, not the multiple the word "flash" implies [5]. Add retries and the gap closes: if the cheap model fails one call in five and you re-run those on the expensive one, you pay both, which the author puts at $18.30 + 20% x $23.00 = $22.90 [15]. That is 0.4% under just running Sonnet for everything, with worse latency attached [6].

The measurement is not hard. Any OpenAI-compatible response returns prompt_tokens and completion_tokens in its usage object; add them up over a few hundred real requests before comparing anything [16]. The author's rules follow from the arithmetic: if input dominates at 10:1 or more, rank on input price and ignore the headline output number; if output dominates below 2:1, rank on output price and look for a low output-to-input multiple [17]. Price the next tier up on your own mix, because if it lands under roughly 1.5x and removes a retry it is cheaper [18]. Do this per endpoint, not per app, since a classifier and a code generator are different workloads [19].

One disclosure the author makes himself: he works on altrouter.ai, which bills the same vendor models under list, Sonnet 5 output at $8.50 against $10.00 and Grok 4.3 at $2.12 against $2.50 [20]. That is a 15% cut on both, which moves the crossover points without changing the method [7].

What to watch: the caveats are load-bearing. This is arithmetic about price, not quality, and it ignores latency, rate limits and prompt caching, any of which can move the answer more than the price gap [21]. Measure your ratio per endpoint, then re-measure after prompt changes, because a bigger system prompt or a longer history window walks your traffic across the crossover without anyone editing a config file [1][16].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories