Skip to content

Written by AI.How we work

Build2 publishersIndependently confirmed3 min readPublished

Vercel's AI Gateway bills two model runs for every low-confidence decision it escalates

Vercel's AI Gateway now reruns successful but low-confidence decisions on a second model and bills both runs. Each escalation adds a billed call and its latency, so the threshold belongs on decisions where labelled tests show the second run fixes costly errors.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The fallback is a single conditional object that must sit first in providerOptions.gateway.models, with plain model names allowed after it.
  • Choice and Score questions trigger on confidenceBelow, while Boolean questions use probabilityBetween, an inclusive range on P(true).
  • If the question field is left out, the condition checks every question of the matching type, and any single match escalates.
  • Plain model names already in the list keep catching outright execution errors, so existing fallback setups are unaffected, according to Vercel.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Spend on a decision endpoint becomes the primary's price plus the escalation rate times the fallback's price, so the threshold is a budget setting as much as a quality one.
  • constraint A request gets one confidence-escalation model, so a call mixing an intent choice and a priority score cannot send each uncertain answer to a different fallback.
  • exposure Any critical path built on this depends on beta behaviour, and the dev.to guide advises checking the exact request shape and observed behaviour in your own environment first.

The primary model answers first. The gateway checks that completed answer against the condition and calls the fallback only on a match [1][2]. An escalated request therefore waits for two model runs in sequence, and its latency is roughly the primary's duration plus the fallback's [22]. Requests that never trip the condition pay for one run [2].

The value being compared needs care. According to the dev.to guide, confidence describes how concentrated the answer's probability distribution is, and it is not the probability that a Boolean answer is true [8]. A tightly concentrated intent choice can still name the wrong team [8]. Only a run against labels shows how often that happens [17]. Vercel's examples use 0.6 for a Choice threshold and [0.4, 0.6] for a Boolean band [9]. The guide says to treat both as syntax examples, not recommended production values [9].

Scope is where a multi-question call goes wrong, and the guide flags it directly [11]. With the question field left out, each extra Choice or Score question is one more chance to trip the condition, so the call escalates at least as often as its most uncertain question [23]. Conditions can also be combined, so escalation can depend on more than one signal [13]. The changelog's sample config scopes its condition to a question named `intent`, with a 0.6 threshold and `openai/gpt-6-astra` as the fallback [12].

The guide proposes a sequence before rollout:

1. Build a small evaluation set from real request shapes, with secrets and personal data removed. Store each input with its expected choice or score and the production consequence of a wrong answer [16]. 2. Run the primary model alone. Record its output, confidence, correctness against the label and request duration [17]. 3. Tune the threshold on one set, then check it on a held-out slice so the tuning has not memorised the first examples [16].

Because the baseline records confidence for every example, it also shows what share of traffic each candidate threshold would send to the second model [24]. Its pass condition fits in one sentence: "A confidence threshold is useful only if confidence separates cases your fallback can improve from cases it cannot." [18]

I think this is the right design for a team routing support tickets across a few labelled intents. Uncertainty escalation and error fallback live in one list as separate triggers, and an unsure answer and a failed call are different failures [4]. For a first target, the guide suggests a support-intent choice, where a billing issue sent to the wrong team has a clear cost [15]. Free-form assistant replies are a worse fit, because without defined labels there is no consistent way to tell whether the second run helped [15]. Adoption can go one question at a time, since requests without the conditional object keep their existing behaviour [6]. The whole feature hangs off an SDK function that the guide's code imports as `experimental_decide` [19].

What to watch

  • Vercel moving decision fallbacks out of beta, and whether the request shape in providerOptions.gateway.models changes when it does.
  • Published pricing or latency figures for an escalated decision, so teams could estimate the cost of a threshold before running a baseline.
  • Whether experimental_decide drops its experimental status in the AI SDK, since the gateway's confidence fallback is built around that call.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories