Build1 publisher3 min readPublished
The rate card is not your bill: budget AI features from the workload, not the price list
A worked example on dev.to shows a prompt that grew from 1,500 to 5,000 tokens more than doubling a monthly bill, with no new users and no price change.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A basic AI API cost calculation is: input cost = monthly input tokens / 1,000,000 x input price per million; output cost = monthly output tokens / 1,000,000 x output price per million; monthly API cost = input cost + output cost. The formula is simple but estimating the numbers that go into it is not.
- A single AI API request can cost less than a cent and still turn into a four-figure monthly bill; the price per million tokens is only one variable.
- Example feature: 1,500 input tokens per request, 300 output tokens per request, 10,000 requests per day, 30 active days per month.
- That workload equals 450 million input tokens and 90 million output tokens per month.
- Illustrative model prices used in the example: $1.50 per million input tokens and $6.00 per million output tokens.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A dev.to explainer on estimating AI API costs works through one arithmetic example and lands on a point that most vendor conversations skip: the price per million tokens is one input to the formula, and rarely the one that moves the total [1][17]. If you are signing off on an AI feature's budget, the leverage sits in your prompt size and your call graph, not in the rate card.
Start with the author's base case. A feature using 1,500 input tokens and 300 output tokens per request, at 10,000 requests a day across 30 days, consumes 450 million input tokens and 90 million output tokens a month [2][3]. At an illustrative $1.50 per million input and $6.00 per million output, that is $675 plus $540, or $1,215 a month [4][5]. The unit economics look trivial: about $0.00405 per request, less than half a cent [6]. Three hundred thousand of them are not a rounding error [7].
Notice what the split already tells you. Output is one sixth of the tokens but 44 percent of the cost, because output is priced at four times input in this example [3][4]. The author's point is that two applications with identical total token counts can have different economics, and that the useful question is how many input and output tokens a typical completed task consumes [10][12]. A classifier sends a lot of context and returns a few tokens; a writing assistant generates hundreds or thousands every run; a coding agent reads context repeatedly and produces substantial output across several calls [11].
Then the quiet multiplier. Hold volume and prices constant and let the average prompt grow from 1,500 to 5,000 input tokens, as system instructions, conversation history, retrieved documents, tool descriptions, examples and metadata accumulate [8][9]. Monthly input tokens go to 1.5 billion, input cost to $2,250, and the total to $2,790 [8]. That is 2.3 times the original bill, up $1,575, with no new users and no repricing [8][1]. Input grew 3.33 times but the bill only 2.3 times, because the $540 of output cost did not move [2]. Context management is a cost line, not just a latency and quality problem [16].
Fan-out is larger still. One button press can trigger a classification, a retrieval, a reasoning call, a second model call with the retrieved context, an evaluation step and a retry [13]. The author's caution is that if one user action generates six model requests, estimating from user actions alone dramatically underestimates usage [14]. Applied to the base example, six calls of that shape is $7,290 a month rather than $1,215 [5]. Which is why the unit to budget in is cost per completed workflow, not cost per message [15].
Put the rate card in perspective. Halving both prices in the base case saves $607.50 a month [6]. One unbudgeted second call in the chain adds $1,215 [7]. A 100 percent discount cannot rescue you from a call graph you have not counted.
What to watch: instrument tokens in and out per completed workflow, not per request, and alarm on the ratio of model calls to user actions. If that ratio drifts, including retries, your bill moves before anyone renegotiates a contract [13][14].