Skip to content

Build1 publisher3 min readPublished

Your inference bill is an architecture defect: declare the task before you call the model

A dev.to walkthrough argues that routing every job through one 'strongest model' helper hides the economics. The cheap fix is declaring task class and requirements before the request leaves your code.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • An AI product can become expensive without doing anything obviously wrong: the prompts work, the model answers correctly, users are getting value, then usage grows and the inference bill grows much faster than expected.
  • One common reason for runaway inference cost is architectural: every task is being sent through roughly the same model path.
  • A document extraction step uses the same model as a difficult reasoning task.
  • A simple classification gets the same reasoning effort as a complex investigation.
  • A repeated 20,000-token workspace context gets sent again and again.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A walkthrough published on dev.to makes a claim worth taking seriously: an AI product can become expensive without anything obviously going wrong, because the prompts work, the model answers correctly, users get value, and then usage grows and the inference bill grows much faster than expected [1]. Its diagnosis is not pricing but structure, namely that every task is being sent through roughly the same model path [2].

The symptoms it lists are recognisable. A document extraction step uses the same model as a difficult reasoning task [3]. A simple classification gets the same reasoning effort as a complex investigation [4]. A repeated 20,000-token workspace context is sent again and again [5]. A model receives hundreds of tool results just to filter and sort them [6]. None of that is broken; the workflow is spending expensive model intelligence on work that does not always need it [7].

The mechanism by which this happens is a single helper. The weak design is a call to `ai.generate` with `model: "strongest-model"`, which every feature eventually calls [8]. According to the piece, that pattern is easy to build and it also hides the economics [9]. The proposed replacement is not a new provider or a discount: it is declaring, in the application, what the task requires before the request reaches the model [10]. The example enumerates seven task names, from extract and classify through research and complex_agent [11], which makes the model choice a consequence of the job rather than a hard-coded default [12].

The task name alone is not enough, since two extraction tasks may have very different requirements and a short invoice and a 200-page legal document should not necessarily follow the same path [13]. Hence a profile carrying six declared attributes: task, complexity, latency, reasoning, volume, and a deterministic post-processing flag [14]. Downstream, the advice is restraint: do not start with twenty routing combinations, because three tiers - economy, balanced, frontier - are enough for many products [15]. Economy covers high-volume work with predictable output such as extraction, classification and tagging [16]; balanced covers support responses, summaries and customer-facing assistants [17]; frontier is reserved for work where better judgment materially changes the outcome, such as ambiguous research or complex agent orchestration [18].

The routing rule itself is deliberately dumb. No machine learning; a rules version is easier to inspect [19], and the sample function returns frontier when complexity or reasoning is high, balanced when either is medium, and economy otherwise [20]. Note the gap: only two of the six declared profile fields actually decide the tier, so latency, volume and the post-processing flag are collected but unused by the sample rule [21]. That is a defensible starting point, not a finished design. Two further structural notes carry weight: keep product logic separate from provider configuration, so a pricing or performance change is a mapping edit rather than a rewrite of every feature [22]; and route reasoning effort independently, because high effort on "extract the invoice number, customer name, and total" adds cost without product value, while comparing five contracts for conflicting obligations may justify it [23].

What to watch is what the excerpt does not contain: no prices, no benchmark, no measured saving from any deployment, and the text breaks off mid-way through the reasoning-selection function [24]. Treat the tier boundaries as a hypothesis your own logs must settle. The testable part is cheap - log the declared profile alongside spend per call, then check how much of your frontier traffic was declared low-complexity by the feature that requested it.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories