Skip to content

Build1 publisher3 min readPublished

Mystery model Union Alpha hit a billion tokens a minute before vanishing from listings and being revealed as Pareto

OpenRouter has dropped the stealth listing. The model is Pareto 26.9, from unbiased.ai, at $2.50 per million input tokens before a 10 October launch. That is double what 26.8 charged.

The Engineer · Build desk

Illustration accompanying Mystery model Union Alpha hit a billion tokens a minute before vanishing from listings and being revealed as Pareto

What happened

  • A search of the OpenRouter API for "union-alpha" on 18 September returned zero results, and the id stealth/union-alpha was no longer in the listing.
  • The model was identified on 17 September as Pareto, built by unbiased.ai, which had shipped it unnamed to see how real users reacted before an official launch set for 10 October 2026.
  • unbiased.ai said demand reached a billion tokens a minute within a day, speed fell to unusable levels, and AWS tripling its compute overnight was still not enough.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Anyone who pointed code at the free stealth endpoint is picking a replacement this week, and the 10 October launch will be quoted against the 26.9 sheet rather than the free tier they tested on.
  • cost Input-heavy callers absorb most of the version change: a request with ten input tokens per output token costs 73 percent more on 26.9 than on 26.8.
  • constraint A buyer cannot reproduce a Pareto result by pinning a model version, because the blend membership is unpublished; the only handle on offer is advance notice of changes.
  • precedent Citing a vendor pricing page in a design doc now comes with an expiry date, since the 26.8 announcement went blank while the 26.9 numbers were still being checked.

The argument worth reading is the one against routing. A router reads a prompt, scores its complexity and picks one model [8]. unbiased.ai says it is not doing that, and its explainer, in Nokka's Thai rendering, says it would already be doing so if that could be done reliably [8]. The obstacle it names is the cache: swap models mid-conversation and the prompt cache goes with them, and at scale the misses eat the whole saving [9]. Cloudflare's description of Union Alpha matches unbiased.ai's description of Pareto, several models called in parallel on every request and synthesized into one answer over a single API for text and images [5][4]. The wording here is translated from a Thai-language post drafted with deepseek-v4.1-flash through Hermes Agent and edited by Nokka [25].

That design puts a floor under the vendor's cost. Every served request pays for several model calls, and the customer is billed for one answer [4]. The published cached-input rate of $0.25 per million tokens only pays off if the same blend members keep seeing the same prefixes [13]. Prefix stability is what a router gives up [9]. Nokka reports Pareto's per-token price is still the lowest in the competitor table he pulled from prices.json [21].

His first draft reported a conflict, $1.25/$6.25 against $2.50/$7.50, and said he could not tell which was live; on recheck he had misread two versions as two sources, and he left the error in the published text on purpose [14][24]. Pareto 26.8 launched at $1.25 input, $0.15 cached and $6.25 output per million tokens [12]. The version serving now is 26.9 at $2.50, $0.25 and $7.50 [13][15]. He confirmed the current figures in four places that agree, among them /how/, /model-card/, prices.json checked on 17 September, and the OpenRouter API [19][17][18]. The /pricing/ page changed while he was editing, from the 26.8 numbers with a competitor note dated 1 August to the 26.9 numbers dated 17 September [16]. The page that announced 26.8 now returns nothing but its title, so the launch prices have no verifiable source left [20].

Input doubled and output rose by a fifth [1]. For a request of 10,000 input tokens and 1,000 output tokens, the bill moves from $0.01875 to $0.0325, up 73 percent [2]. Retrieval and long-context callers take close to the full doubling; short-prompt chat takes close to the 20 percent [1].

The benchmark line is the part I would not carry into a plan. That first draft cited 86.0 on SWE-Bench, 90.4 on GPQA, 77.9 on MMMU-Pro and a 73.5 average [22]. Those numbers describe a blend whose membership unbiased.ai declines to publish, on the grounds that a published list would be out of date within a few weeks [10]. What the company offers instead is procedural, a commitment to announce routing or membership changes before they go into service [11]. For a September score to hold in October, the blend would have to be the same blend. unbiased.ai says it tests new models as they ship and drops the ones it no longer wants in the blend [10].

What to watch

  • Whether the 10 October launch keeps the 26.9 price sheet or arrives with another version number and another rate.
  • Whether unbiased.ai publishes its first routing or blend-membership change notice, the one auditable commitment it has made.
  • Whether a named Pareto listing reappears on OpenRouter, and at what input and output rates.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories