Build1 distinct publisher3 min readPublished
On a dev.to walkthrough's illustrative support assistant, the tier costing 4.9 times more per token ends up at under half the monthly cost once the humans who clean up its failures are priced in. The whole calculation is four multiplications.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Work out where the 0.0026 in the mid tier's retry line comes from. It is 156.00 divided by 60,000 conversations, the tier's own blended cost per call [2]. Retries are therefore billed at the average, which is a fair simplification and also why the retry term vanishes: 6,000 second attempts come to 15.60, against 7,200.00 of escalated human time [10].
Then compare the two deltas. Going from mid to frontier adds 609.00 of token spend and removes 4,800.00 of escalation cost, so the human term moves about 7.9 times more money than the rate card does [3].
That ratio is why the escalation rate is the only line worth arguing about, and you can solve for the point where it stops mattering. Hold the frontier tier at 1 percent and its monthly total at 3,165.00 [11]. The mid tier's non-escalation costs are 156.00 plus 15.60 [10]. That leaves 2,993.40 of escalation budget, and at 4.00 per escalation across 60,000 conversations the mid tier breaks even at 1.25 percent [4]. Its winning window is a quarter of a percentage point wide [5]. Measuring an escalation rate to that resolution means labelling real transcripts, which is more work than the four multiplications and is the work that actually picks the tier.
The table is illustrative and says so [3]. For its ranking to carry into your system, four things have to hold: an input to output ratio near 4.3 to 1 [6], a two point escalation gap between the tiers [10], six-minute human handling, and a 40.00 loaded support hour [9]. The ratio is doing more than it looks. The frontier input rate is 6.25x the mid tier while its output rate is only 3.75x [8], and the blended multiple that falls out, 765.00 over 156.00, is 4.9x [7], sitting near the input multiple because this workload reads about four times as much as it writes [6]. Push the mix toward long generations and the blended multiple slides toward 3.75x [11], and the mid tier keeps more of the token advantage it started with.
The input count is where estimates rot. A 40 token user message with a 900 token system prompt and three retrieved documents in front of it is billed in full on every call [14]. Pull the transcript out of your logs and count everything you send.
The article declines to print a single real rate, on the grounds that published prices move and a number baked into an article goes stale silently [12]. That is the right call for a method piece, and it means the inversion in the table is a property of the placeholders rather than of anyone's current price list [5]. The author also refuses the easy summary that the frontier tier always wins, and names the escalation rate as the deciding variable instead [13]. A cent per conversation has never blocked a launch [7]; it is review meetings, not cents, that decide what gets contested.
Four multiplications produce a number a product manager can act on. Getting the decision itself right still takes a day of transcript labelling.
Ranked by verification strength, evidence, and original report placement.
The method prices a workload with four numbers multiplied before any code is written: calls per day, input tokens per call including the system prompt, retrieved context and conversation history resent every turn, output tokens per call, and days per month actually run.
The worked example is an illustrative support assistant handling 2,000 conversations a day, averaging 1,500 input tokens and 350 output tokens per conversation, run 30 days a month; the author states these are illustrative numbers, not measurements, and says to replace every one of them.
Monthly volumes for that workload: 90.0M input tokens, 21.0M output tokens, and 60,000 conversations.
The article uses placeholder rates, explicitly not any model's real price: frontier tier at 5.00 per million input tokens and 15.00 per million output, mid tier at 0.80 per million input and 4.00 per million output.
At those placeholder rates the frontier tier totals 450.00 input plus 315.00 output for 765.00 a month, and the mid tier totals 72.00 plus 84.00 for 156.00 a month.
The difference is about one cent per call, which the author says will never survive a design review as an objection, but 609.00 a month at 60,000 conversations and a headcount at ten times the volume.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Three manual interventions in a month, and every guard was working as designed1 distinct publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 distinct publisher
build
Force the tool call, then hand Lightsail a long-lived key1 distinct publisher
build
Separating moderation rejections moved one API gateway's success rate from 95.5% to 98.9%1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Sums check, inputs invented
Every multiplication reproduces: 90.0M tokens at 5.00 is 450.00, the 0.0026 retry unit is simply 156.00 spread across 60,000 conversations, and the 7,371.60 and 3,165.00 totals follow. What the arithmetic cannot supply is the world — the prices are declared placeholders and the 3-versus-1 percent escalation gap that drives the whole result is a supposition, so the piece is internally airtight and externally unverified.
Nothing shipped to observe
A costing method leaves no footprint we can count. This reporting names no team that ran the four multiplications, produces no logged escalation rates from live traffic, and identifies no model whose price you could even substitute in — so there is no uptake to measure, only a spreadsheet argument.
Quotable beyond its evidence
'7.9 times more money' reads like a measurement and is an artefact: it divides the gap between two escalation rates the author chose by a rate difference he also chose. The overstatement risk lives in how travel-ready that number is, not in his prose — he refuses the 'always use the frontier model' reading, calls the rates assumptions, and tells you to instrument resolution rate first. Strip the invented inputs and one claim still stands up: if human cleanup dwarfs token spend, the token price is the wrong thing to argue about.
One link, no vendor to please
The only thing being sold is a pointer to a linked calculator for current dated prices. Otherwise the pressures run the other way: no provider is named, no rate is quoted, and the conclusion — spend more on tokens, measure your humans — flatters nobody's price list. Declining to print real prices is genuinely good practice and also keeps the piece from ageing, which is convenient for its author.
Confident in the sums, not the world
We can stand behind the arithmetic and the structural point without reservation, because both are checkable line by line on the page. We cannot stand behind any dollar figure surviving contact with a real workload: one publisher, one invented example, zero measured escalation rates, and a conclusion that flips on a quarter of a percentage point.