Build1 publisher3 min readPublished
Cloudflare's Auto Router moves model selection from the user into AI Gateway
Cloudflare's Auto Router, now in public beta, picks a model per AI Gateway request and cut costs by up to 30% in the company's own OpenCode use. How much of that transfers depends on how much of a team's traffic never needed a frontier model.
The Engineer · Build desk

What happened
- Teams opt in by setting the model name to cloudflare/auto, and the gateway then chooses a model it judges capable enough for each task.
- Harnesses such as OpenCode, Claude Code and Codex leave model selection to each user, who often ends up on a model that is overkill for the job.
- On Cloudflare's internal knowledge-work benchmark, Auto Router performed similarly to other leading daily-driver models at 80% of GPT-6 Sol's cost and 35% of Claude Opus 5.5's.
- Cloudflare also runs the router inside Cloudflare OS, its custom agent harness, and reports coding results comparable with frontier models.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Model choice becomes an organisation-level gateway setting, so platform teams decide whether routing applies and individual users stop being the main cost control.
- constraint The 65% saving against Opus holds for a request mix like Cloudflare's office-tool benchmark; a coding-heavy team should plan around the 30% ceiling or less.
- cost Comparing models by price per million tokens can mislead, because a cheaper model can burn more tokens; cost per finished task is the figure to measure before and after switching.
- exposure Once users stop picking models, a drop in answer quality is harder to attribute, so a bad routing decision has to be caught in the gateway's own records.
AI Gateway can make the choice because every request from every user, agent and tool already passes through it, according to Cloudflare [5]. The company's earlier controls were budgets, spend limits and identity-aware analytics. Those still depend on each person making a cost-conscious choice on every request [6]. The router moves that decision into the gateway, and Cloudflare says users keep access to the most capable models when their work needs them [7].
Cloudflare's example is an email summary. It says that job does not need Opus-level intelligence, and that it still would not block Opus outright for its security engineering team [4]. A per-request router lets both hold. The summary goes to a cheaper model, and the frontier model stays available for the work that needs it [7].
The post gives two savings figures from two measurements. The up-to-30% comes from internal OpenCode use, measured against frontier-only use [2]. The benchmark cost ratios translate to a 20% saving against Sol and a 65% saving against Opus [1][2]. The internal ceiling sits between them [3].
Both figures come from Cloudflare's own workload. The benchmark is 97 tasks run three times per model on simulated office tools, and each task has to end in a verifiable answer or a completed action [10][11]. Cloudflare says savings grow with the share of work that never needed frontier rates [12]. It also says the router does best on the broad knowledge work of a large organisation with technical and non-technical teams [13]. For the 65% figure to transfer, a team's traffic has to resemble that benchmark's office-tool tasks. I'd expect a team that mostly sends hard coding work to land nearer the OpenCode ceiling, or under it [2][11].
The statistics are done with care. Cloudflare's 95% confidence intervals come from 10,000 bootstrap resamples at the task level, keeping all three repetitions of a task together [10]. Resampling single runs would treat three attempts at one task as independent evidence and make the intervals look tighter than the data supports [10].
The strongest point in the post is about price. A model that looks cheaper per token can use disproportionately more tokens to solve the same problem [14]. "A router should minimize predicted trajectory cost, not just load-balance by dollars per million tokens," Cloudflare wrote [15]. That is the right objective for anyone building a router. One tuned to the price sheet would favour the cheapest model per million tokens and could still pay more per finished task [14].
The post opens on the line "The best savings are the ones users never notice" [16]. Users will not notice a routing mistake either. Before turning on cloudflare/auto for a whole organisation [1], I would want the gateway to record which model answered each request, so a weak answer can be traced back to the router's choice.
What to watch
- Pricing and terms for Auto Router when it leaves public beta.
- Customer-run comparisons on coding-heavy traffic, where Cloudflare's own figures put the savings case at its weakest.
- Whether AI Gateway lets administrators pin a team such as security engineering to a frontier model while routing everyone else.