Skip to content

Product1 publisher3 min readPublished

Routing easy jobs to cheaper models got Uber more than nine times the AI usage

A Fast Company column reports that many of Anthropic's US customers are picking cheaper models over its flagship, and it uses Uber's token bill as the template for how buyers now split work by price.

The Product Desk · Product desk

Illustration accompanying Routing easy jobs to cheaper models got Uber more than nine times the AI usage

What happened

  • Uber consumed a year's worth of AI tokens in just four months earlier this year, according to the column.
  • Uber then sent simple jobs to cheaper models and kept the most expensive systems for the hardest work, and AI usage across the business rose more than ninefold without a corresponding rise in spending.
  • The column also reports that Anthropic and OpenAI have both posted sharply accelerating revenue in recent months, with Anthropic recording its first adjusted operating profit.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • contradiction A finance lead can quote the same column on either side, because the accelerating-revenue line sits next to a downgrade claim the column never sizes or ties to a named customer.
  • decision A team that planned to settle on one vendor now has a per-workload choice to defend instead, and each cheap-model assignment needs an eval before anyone can sign off on it.
  • cost Portfolio buying moves cost off the invoice and onto the team: prompts, evals and fallbacks per model, maintained by whoever owns the router.
  • precedent If labs start pricing tiers the way the column urges, the flagship becomes the on-ramp and the paid product becomes the layer that picks which model runs.

Routing costs something before it saves anything. To send a job to a cheaper model you have to know the cheaper model does that job, which means an eval per task and a fallback for when it fails. Someone owns that eval the next time a vendor ships a new version. Uber's payoff, as the Fast Company column tells it, came only after it stopped treating every task the same [6].

The pressure that produced the change shows up in the token bill. A year's worth of tokens in four months is three times the rate the budget assumed [5][1].

The decision looks like a vendor choice: pick the best model and standardize on it. What the buyer in the column did was sort work by difficulty and pay frontier prices only for the hard part [6]. The column's wider claim is that companies will stop standardizing on one frontier model and start buying outcomes, at which point the model becomes an input to someone else's product [9].

On Anthropic, the column asserts more than it shows. It says many US customers are choosing cheaper models over the most advanced option, and that the company is finding it harder to persuade them to pay for its leading model, which the column calls Fable 5 [1][4]. It does not quantify how many moved or identify a single customer [12]. Both labs have also put cheaper models below the frontier while holding the price of their flagships [8].

For the person who has to write the routing policy on Monday, two axes do most of the sorting. One is whether a human reads the output before it leaves the team. The other is whether a wrong answer costs money you can put a figure on. Reviewed output with cheap errors goes to the cheapest model that passes your eval. Unreviewed output with expensive errors keeps the flagship until an eval says otherwise. Where a person checks the work but a mistake is costly, run the cheap model first and pay for the second look. Cheap errors with no reviewer get the cheap model and sampled audits.

I would set the default at the cheapest model that clears your eval, with a written reason on file for every workload that uses the flagship. The tradeoff is that you now maintain an eval suite and a router, and every model upgrade reopens both. The column's own prescription is aimed at the labs: sell portfolios at several price points, compete on the outcome, and let today's frontier model become tomorrow's cheap way in [10]. It expects the contest to move to the products and services wrapped around the models [13].

What to watch

  • Whether Anthropic's IPO disclosures break out revenue by flagship versus cheaper models. That break-out is the only way the downgrade claim gets a number.
  • Whether either lab ships a priced portfolio or an automatic task router as a product, rather than separate models on separate price cards.
  • Whether a second named enterprise publishes routing savings with figures attached, so Uber stops being the only data point.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories