Skip to content

Build1 publisher3 min readPublished

Open-weight models took 78.4% of Vercel gateway tokens on a single September day

Guillermo Rauch's daily snapshot counts tokens inside one managed gateway's self-selected traffic. Vercel's monthly index carries the figure that prices the routing decision, 56% of tokens against 14% of estimated spend.

The Engineer · Build desk

Photograph accompanying Open-weight models took 78.4% of Vercel gateway tokens on a single September day
Photo: rauchg.com

What happened

  • Vercel founder Guillermo Rauch said open-weight models handled 78.4% of token volume through the company's AI Gateway on a September 19th daily snapshot, against 21.6% for closed models.
  • Vercel's September production index, which covers traffic through August, put open-weight models at 56% of gateway tokens while accounting for 14% of estimated spending.
  • Rauch said Moonshot AI and DeepSeek ranked third and fourth by estimated spend that day, and that their combined spending with Z.ai exceeded OpenAI's share on the gateway.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost On the August split, the average closed-model token cost roughly 7.8 times the average open-weight one. That multiple is the size of the saving a routing layer can capture on this traffic mix.
  • constraint Anyone copying the percentage has to check their own token composition first, because cached-input and cache-creation tokens count toward share and a caching-heavy open-model workload inflates volume without a matching bill.
  • decision With rules applied at the credential level, swapping or banning a model is a gateway configuration change, so the choice moves from application teams to whoever holds the gateway credentials.

Token share and spend share do not measure the same thing. Vercel counts input, output, reasoning, cached-input and cache-creation tokens in the share, and says long contexts, automated agents and a small number of computationally heavy workloads can push it up [8]. Spend is computed at list prices, and it leaves out request counts and the business value of the work produced [9]. One agent loop replaying a long context can therefore outweigh a lot of short interactive requests in the count.

Run the August pair through the same split and the price gap appears. Open-weight models took 56% of tokens for 14% of estimated spend, which leaves closed models 44% of tokens and 86% of spend [5][19]. The average closed-model token on that gateway cost about 7.8 times the average open-weight token [20]. Vercel said the shift helped cut the gateway's average price per token by 23.2% during August [7].

For that ratio to describe another team's bill, its traffic has to be shaped like the gateway's. The August comparison suggests developers are assigning large volumes of work to cheaper open models while reserving expensive frontier models for tasks where they believe the premium is justified [12]. The team also has to be buying inference at list prices through providers [9]. And the sample has to be comparable. The daily number is weakest here: 78.4% counts tokens inside one managed gateway's self-selected customer traffic on one day, not three-quarters of the model market [10]. Vercel says the gateway routes tens of trillions of tokens a month [11]. That gives the sample scale. Its scope is one gateway's customers, not all AI usage.

The part that makes a swap cheap is the control surface. The gateway reached general availability on August 21st, 2025, with one API for hundreds of models, usage analytics, automatic failover and routing based on cost, latency or availability [13]. Vercel says it adds no markup to token prices [14]. In July the company added gateway-level routing rules that let teams rewrite requests to a different model, or block an unapproved model, across every application using their credentials [15]. Enforcing that at the credential boundary means one change covers services whose code nobody wants to reopen, and a model ban works whether or not every repository honours it.

Aligned News highlighted the snapshot on September 20th, after Rauch published the underlying data [16]. Rauch attached a qualification to the spend ranking: the dollars are inference bought through providers, mostly in the United States, and should not be read as revenue flowing directly to Moonshot AI, DeepSeek or Z.ai [4]. He called the 78.4% a possible record day for open models on the gateway [2]. runtimewire notes that Vercel benefits as models become easier to swap, and that a model lab wants developers to stay inside its own family [17][18]. The company selling the routing layer is also the company counting the tokens. That does not make the count wrong. I read the denominators carefully.

The trend line is the part I would plan against. Open-weight token share on the gateway was 7% in December 2025 and 56% in the index covering August, a move of 49 points [6][5][22]. The daily figure sits 22.4 points above that monthly one [21]. In my view a budget built on the daily figure is priced off a peak.

What to watch

  • Whether the next monthly production index moves the 56% token share toward the daily 78.4%, and whether the 14% spend share moves with it.
  • Whether Vercel keeps publishing the spend ranking now that Moonshot AI, DeepSeek and Z.ai outrank OpenAI on it.
  • Whether the no-markup pricing holds as the gateway's average price per token keeps falling.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories