Skip to content

Build4 publishers3 min readPublished

Three frontier launches in a day, all pitched on price. Open weights set the ceiling.

Grok 4.6, Qwen3.8-Max and DeepSeek V4-Pro shipped inside about 24 hours, and two of the three came with downloadable weights. The benchmarks existed to justify a cheaper invoice.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Three frontier launches in a day, all pitched on price. Open weights set the ceiling.
Generated illustration

What happened

  • Grok 4.6 launched Wednesday, Qwen 3.8-Max a few hours later, and DeepSeek V4-Pro on Thursday: three frontier models in about 24 hours, all pitched on cost because capabilities are assumed.
  • Two of the three frontier launches came with downloadable weights, which is why the ceiling on what a closed lab can charge is increasingly set by companies giving their models away.
  • SpaceXAI announced Grok 4.6 in two sentences, describing frontier intelligence and a "significant improvement over Grok 4.5 at the same price."
  • xAI reports Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 High and level with GPT-5.6 Sol Max, while Fable 5 Max scored 62.
  • Artificial Analysis placed Claude Opus 5 at 63 and Claude Fable 5 at 62 on its Intelligence Index under the configurations it evaluated.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Three frontier models shipped in roughly 24 hours this week, and per The New Stack all three were sold on cost rather than capability: Grok 4.6 on Wednesday, Qwen3.8-Max a few hours later, DeepSeek V4-Pro on Thursday [1]. Two of the three came with downloadable weights, which is the part that matters for anyone signing an inference contract: the ceiling on what a closed lab can charge is increasingly set by companies giving the model away [2].

Start with the one that did not ship weights. xAI announced Grok 4.6 as frontier intelligence and a "significant improvement over Grok 4.5 at the same price" [3]. Artificial Analysis scored it 61 on its Intelligence Index, up from 56 for Grok 4.5 High and level with GPT-5.6 Sol Max, with Claude Fable 5 at 62 and Claude Opus 5 at 63 [4][5]. API pricing stayed at $2 per million input tokens and $6 per million output, with the fast variant at double [6], against $5/$25 for Opus 5 and $5/$30 for GPT-5.6 Sol [7]. That is 60 percent less on input and 80 percent less on output than GPT-5.6 Sol [8]. Investor Gavin Baker put Grok 4.6 at 80 percent cheaper on input and 88 percent cheaper on output than Fable 5 Max, and called it "Pareto dominant" [9].

The reason the price did not move is that the model did not get bigger. Product analyst Aakash Gupta says Grok 4.6 runs on the same 1.5 trillion parameters as 4.5, with the gains coming out of post-training, including a 66 percent jump on Terminal-Bench in one release [10]. Same weights, same serving cost, same bill.

The scoreboard is not uniform, and the sources disagree on the headline. The New Stack has Grok 4.6 topping GDPVal-AA at 1,753 Elo, just past Fable 5 Max at 1,741 [11]; The Decoder puts it second on GDPval-AA v2 at the same Elo, behind Claude Opus 5 [12]. xAI's own listing has it leading on GDPVal-AA v2 and AA-Briefcase while trailing GPT-5.6 Sol Max and Fable 5 Max on DeepSWE and Terminal-Bench [13]. On CursorBench v3.2 it scored 69.9 percent against 70.5 percent for Fable 5 Max [14].

The open side is what disciplines that pricing. When Alibaba announced Qwen3.8-Max, consultant Jeff Brokaw dismissed an open-weights promise with no release as "the API business model wearing an open source jacket for the launch photo"; the weights are now out [15]. DeepSeek V4-Pro followed with native support for OpenAI's Response API and a two-tier price table, off-peak at half of peak [16]. If the fallback is a file on Hugging Face, a lab ships the better model at the old price and absorbs the difference.

Box CEO Aaron Levie's read is that cheaper tokens do not shrink budgets, they fund work firms already wanted, and that the value moves to the layer that routes and optimises per task [17]. Nvidia shipped exactly that this week: Nemotron 3.5 Lightning, an open 30-billion-parameter model, plus NeMo Switchyard, an open source router; Nvidia's own figures have a Switchyard setup pairing open models with Anthropic's Opus 4.8 at about a third of the cost with frontier accuracy intact [18].

Watch effective cost per task, not the token sheet. Artificial Analysis estimates $0.84 per task for Grok 4.6 and about 53 turns and 0.5 billion input tokens on AA-Briefcase, against roughly 103 turns and 2.0 billion for Opus 5 [19] - four times the input tokens for the same job [20]. The same evaluator notes caching, retries and tool-call overhead usually dominate agent bills [21]. Also watch whether the next closed release dares to raise its price, and whether router savings survive an audit by someone other than the vendor.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories