Skip to content

Project

Vercel AI Gateway

Vercel's unified API gateway that routes calls to multiple AI model providers through one interface, passing through each provider's native pricing.

Current stories

buildConfirmed22 publishers

Prompt length sets what Anthropic's Haiku 5.5 price cut is worth to each workload

Anthropic launched Claude Haiku 5.5 at an average price about 75% below Haiku 4.5. The saving varies widely with prompt length, so teams moving classification, support or query traffic need to price their own requests before they switch models.

Perspective Coverage

22 publishers
Builder
Builder 47%
Operator
Operator 36%
Investor
Investor 17%

Reality

Evidence66
Adoption30
Hype gap+25
Incentives70
Confidence65
buildConfirmed7 publishers

Xiaomi priced six days of reinforcement learning at $3.47 million

Xiaomi's $850,000 and $2.62 million cover 30 reinforcement-learning steps apiece on models it had already pretrained. Xiaomi published the environments and the RL code but not the 7,000-plus task datasets behind them.

Perspective Coverage

7 publishers
Builder
Builder 56%
Operator
Operator 23%
Investor
Investor 21%

Reality

Evidence60
Adoption25
Hype gap+30
Incentives65
Confidence62
buildConfirmed8 publishers

Cursor's task-cost table shrinks Grok 4.7's 7.5x list-price gap to about 1.9x

SpaceXAI lists Grok 4.7 at $2 and $6 per million tokens against $10 and $50 for GPT-6 Astra and Claude Fable 5.1, and the only per-task figures in the record put a Cursor run at $4.69 against about $9.

Perspective Coverage

8 publishers
Builder
Builder 51%
Operator
Operator 32%
Investor
Investor 17%

Reality

Evidence60
Adoption30
Hype gap+25
Incentives65
Confidence58
buildConfirmed6 publishers

DeepSeek's V4.1-Flash report prices a live agent's memory at 890 bytes per token

The quarter-size global KV cache and eighth-size persistent storage are serving-cost claims rather than benchmark scores, and banking the second one requires a prefix-cache tier your stack has to already run.

Perspective Coverage

6 publishers
Builder
Builder 53%
Operator
Operator 33%
Investor
Investor 14%

Reality

Evidence66
Adoption42
Hype gap+14
Incentives64
Confidence63
buildConfirmed8 publishers

Gemini 3.7 Flash's real pitch is fewer dead agent runs, and it is half price until December

Google's new workhorse Flash is a model-string swap for anyone on AI Gateway, discounted through 31 December 2026. The claim worth testing is reduced tool-calling loop failures, not benchmark deltas.

Perspective Coverage

8 publishers
Builder
Builder 54%
Operator
Operator 24%
Investor
Investor 22%

Reality

Evidence55
Adoption45
Hype gap+25
Incentives70
Confidence60
buildConfirmed21 publishers

Mistral's Arthur Mensch claims a cybersecurity edge over unnamed Chinese models

Mistral CEO Arthur Mensch said on October 6th that the company's newest model beats unnamed Chinese rivals on cybersecurity. He gave no tests or scores with the claim, so buyers looking for a supplier outside the US and China cannot yet use it to choose one.

Perspective Coverage

25 publishers
Builder
Builder 38%
Operator
Operator 34%
Investor
Investor 28%

Reality

Evidence60
Adoption20
Hype gap+35
Incentives75
Confidence65
buildConfirmed9 publishers

Nano Banana 2.1's half-price images save money only on input-light jobs

Google priced Nano Banana 2.1 at half Nano Banana 2's per-image rate, $0.0336 for a 1K image, while tripling what it charges for input tokens. Reference-heavy edits can cost more after switching, and no published benchmark backs the quality claims.

Publishers:ai.google.devblog.vercel.comcellcog.aidecrypt.codev.tonokiapoweruser.comnotebookcheck.netruntimewire.comtechrepublic.comthe-decoder.com

Perspective Coverage

10 publishers
Builder
Builder 53%
Operator
Operator 35%
Investor
Investor 12%

Reality

Evidence64
Adoption30
Hype gap+22
Incentives58
Confidence66
buildConfirmed4 publishers

OpenAI splits classification into a Decisions API priced at 10 cents per million input tokens

OpenAI's Decisions API answers yes/no, pick-one and scale questions for $0.10 per million input tokens, with output free. Jev sells the same three question types for text at less than half that, so OpenAI's case rests on image input and compliance terms.

Perspective Coverage

4 publishers
Builder
Builder 59%
Operator
Operator 24%
Investor
Investor 17%

Reality

Evidence62
Adoption20
Hype gap+15
Incentives45
Confidence60
buildConfirmed12 publishers

OpenAI courts builders with a cheaper model and ChatGPT's 1.2 billion weekly users

OpenAI priced GPT-6.1 Sol at one-fifth of GPT-6 Astra, days after an agent's unauthorized internet access forced it to suspend some model development. Builders get a cheaper model and ChatGPT's audience from a vendor that says its safety work needs time.

Perspective Coverage

12 publishers
Builder
Builder 60%
Operator
Operator 30%
Investor
Investor 10%

Reality

Evidence66
Adoption40
Hype gap+10
Incentives70
Confidence68
buildConfirmed7 publishers

Microsoft's MAI speech lineup documents real-time use only on the voice-output side

Microsoft's MAI-Transcribe-2 covers 60 languages with speaker labels and word timestamps, though streaming is not among its documented features. Live voice products get a fast model for replies and still need another way to hear the caller.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 30%
Investor
Investor 17%

Reality

Evidence35
Adoption15
Hype gap−35
Incentives65
Confidence70
buildConfirmed6 publishers

Gemini 3.8's cheap TTS tier still takes line-by-line stage directions

Google shipped two text-to-speech models with the same per-line performance control and put voice creation on only one of them, so the tier decision depends on whether a project has to invent a voice at all.

Perspective Coverage

7 publishers
Builder
Builder 64%
Operator
Operator 23%
Investor
Investor 13%

Reality

Evidence62
Adoption30
Hype gap+25
Incentives58
Confidence66
buildConfirmed14 publishers

One model string moves Vercel AI Gateway traffic to Claude Sonnet 5.5

Vercel's AI Gateway now routes Claude Sonnet 5.5 through a single model ID, according to a dev.to review of the week's releases. The benchmark and cost figures come only from that third-party review, so a team's own tests decide when regulated workloads move.

Perspective Coverage

14 publishers
Builder
Builder 47%
Operator
Operator 30%
Investor
Investor 23%

Reality

Evidence58
Adoption48
Hype gap+35
Incentives62
Confidence60
buildConfirmed3 publishers

Kimi K3 on its cheapest host undercuts Fireworks' Ember-1 despite a 23% cut in reasoning tokens

Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 30%
Investor
Investor 18%

Reality

Evidence55
Adoption30
Hype gap+25
Incentives70
Confidence58

Earlier coverage

  1. GLM-5.3 keeps GLM-5.2's base model and claims 50% more on coding: plan for shorter eval cycles

    Build · August 16, 2026 · Confirmed5 publishers

  2. Google splits transcription in two, and quietly absorbs your cleanup layer

    Build · August 26, 2026 · Confirmed6 publishers

  3. Vercel moves DNS, renewals and membership into a CLI your pipeline already trusts

    Build · August 27, 2026 · Confirmed2 publishers

  4. Gemini 3.8 Flash's introductory price doubles on December 31, 2026

    Build · September 2, 2026 · Confirmed8 publishers

  5. Spark 1.3's index jump lands on the three tests that carry half the score

    Build · September 3, 2026 · Confirmed6 publishers

  6. Multi-turn GPT Image 2.5 editing runs only through OpenAI's Responses API

    Build · September 13, 2026 · Confirmed7 publishers

  7. Who owns the GPU fleet decides whether LLM routing is a library or a gateway

    Build · September 21, 2026 · One report1 publisher

  8. Filling GLM-5.3-Flash's million-token window costs three times its per-task benchmark price

    Build · September 20, 2026 · Confirmed2 publishers

  9. Rounding a judge model's 0.99 into a price tier billed a synthesis as a lookup

    Build · September 20, 2026 · One report1 publisher

  10. Open-weight models took 78.4% of Vercel gateway tokens on a single September day

    Build · September 20, 2026 · One report1 publisher

  11. A strict enum schema on the baselines erased most of Jev's 14x decision-latency lead

    Build · September 19, 2026 · One report1 publisher

  12. Open-weight models ran 56% of Vercel's gateway tokens for 14 cents of every dollar

    Build · September 19, 2026 · Confirmed2 publishers

  13. Vercel's stacked AI Gateway budgets reject a request at whichever cap runs out first

    Build · September 16, 2026 · One report1 publisher

  14. Open weights take 29% of gateway tokens on a twenty-fifth of the dollars

    Invest · August 30, 2026 · One report1 publisher

  15. Enterprise buyers pay Anthropic 63% of API spend for 31% of the tokens

    Invest · August 29, 2026 · One report1 publisher

  16. Experiential Labs bets its open-source router's traces will train cheaper replacements for rented models

    Build · August 27, 2026 · One report1 publisher