Build4 distinct publishers3 min readUpdated
Grok 4.6, Qwen3.8-Max and DeepSeek V4-Pro shipped inside about 24 hours, and two of the three came with downloadable weights. The benchmarks existed to justify a cheaper invoice.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Grok 4.6, Qwen3.8-Max and DeepSeek V4-Pro shipped inside about 24 hours, and two of the three came with downloadable weights. The benchmarks existed to justify a cheaper invoice.
Three frontier models shipped in roughly 24 hours this week, and per The New Stack all three were sold on cost rather than capability: Grok 4.6 on Wednesday, Qwen3.8-Max a few hours later, DeepSeek V4-Pro on Thursday [1]. Two of the three came with downloadable weights, which is the part that matters for anyone signing an inference contract: the ceiling on what a closed lab can charge is increasingly set by companies giving the model away [2].
Start with the one that did not ship weights. xAI announced Grok 4.6 as frontier intelligence and a "significant improvement over Grok 4.5 at the same price" [3]. Artificial Analysis scored it 61 on its Intelligence Index, up from 56 for Grok 4.5 High and level with GPT-5.6 Sol Max, with Claude Fable 5 at 62 and Claude Opus 5 at 63 [4][5]. API pricing stayed at $2 per million input tokens and $6 per million output, with the fast variant at double [6], against $5/$25 for Opus 5 and $5/$30 for GPT-5.6 Sol [7]. That is 60 percent less on input and 80 percent less on output than GPT-5.6 Sol [8]. Investor Gavin Baker put Grok 4.6 at 80 percent cheaper on input and 88 percent cheaper on output than Fable 5 Max, and called it "Pareto dominant" [9].
The reason the price did not move is that the model did not get bigger. Product analyst Aakash Gupta says Grok 4.6 runs on the same 1.5 trillion parameters as 4.5, with the gains coming out of post-training, including a 66 percent jump on Terminal-Bench in one release [10]. Same weights, same serving cost, same bill.
The scoreboard is not uniform, and the sources disagree on the headline. The New Stack has Grok 4.6 topping GDPVal-AA at 1,753 Elo, just past Fable 5 Max at 1,741 [11]; The Decoder puts it second on GDPval-AA v2 at the same Elo, behind Claude Opus 5 [12]. xAI's own listing has it leading on GDPVal-AA v2 and AA-Briefcase while trailing GPT-5.6 Sol Max and Fable 5 Max on DeepSWE and Terminal-Bench [13]. On CursorBench v3.2 it scored 69.9 percent against 70.5 percent for Fable 5 Max [14].
The open side is what disciplines that pricing. When Alibaba announced Qwen3.8-Max, consultant Jeff Brokaw dismissed an open-weights promise with no release as "the API business model wearing an open source jacket for the launch photo"; the weights are now out [15]. DeepSeek V4-Pro followed with native support for OpenAI's Response API and a two-tier price table, off-peak at half of peak [16]. If the fallback is a file on Hugging Face, a lab ships the better model at the old price and absorbs the difference.
Box CEO Aaron Levie's read is that cheaper tokens do not shrink budgets, they fund work firms already wanted, and that the value moves to the layer that routes and optimises per task [17]. Nvidia shipped exactly that this week: Nemotron 3.5 Lightning, an open 30-billion-parameter model, plus NeMo Switchyard, an open source router; Nvidia's own figures have a Switchyard setup pairing open models with Anthropic's Opus 4.8 at about a third of the cost with frontier accuracy intact [18].
Watch effective cost per task, not the token sheet. Artificial Analysis estimates $0.84 per task for Grok 4.6 and about 53 turns and 0.5 billion input tokens on AA-Briefcase, against roughly 103 turns and 2.0 billion for Opus 5 [19] - four times the input tokens for the same job [20]. The same evaluator notes caching, retries and tool-call overhead usually dominate agent bills [21]. Also watch whether the next closed release dares to raise its price, and whether router savings survive an audit by someone other than the vendor.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
xAI reports Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 High and level with GPT-5.6 Sol Max, while Fable 5 Max scored 62.
Artificial Analysis placed Claude Opus 5 at 63 and Claude Fable 5 at 62 on its Intelligence Index under the configurations it evaluated.
Grok 4.6 launched Wednesday, Qwen 3.8-Max a few hours later, and DeepSeek V4-Pro on Thursday: three frontier models in about 24 hours, all pitched on cost because capabilities are assumed.
SpaceXAI announced Grok 4.6 in two sentences, describing frontier intelligence and a "significant improvement over Grok 4.5 at the same price."
Grok 4.6 API pricing is $2 per million input tokens and $6 per million output tokens, with the fast variant costing twice as much.
Grok 4.6 pricing of $2/$6 per million tokens is more than 60 percent cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30).
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Strong on price and evaluator numbers, thin on the market thesis
The load-bearing specifics - $2/$6 pricing, the 61 Intelligence Index score, the 53-versus-103 turn gap, mixed DeepSWE/Terminal-Bench results - appear in three or four independent publishers, two of them citing vendor documentation and an independent evaluator. Weaker links: the leaderboard rank conflicts between outlets, and the parameter count, Terminal-Bench delta, Fable 5 Max discount percentages, DeepSeek pricing table and Nvidia cost ratio each rest on a single publisher or on the vendor's own numbers.
Broad day-one distribution, no usage evidence
Availability is concrete and multi-channel on release day - xAI API, Cursor, Grok Build, OpenRouter, Vercel, Cloudflare, with doubled first-week quotas - and the surrounding week adds Qwen3.8-Max weights, DeepSeek V4-Pro and Nvidia's open router. What is missing is any usage disclosure, customer count, token volume or production deployment: the routing-layer and open-weights-substitution behaviour is argued rather than measured, with the only quantified enterprise split coming second-hand through one outlet.
Discount framing runs ahead of per-job evidence
Headline framings - matched at an 85 percent discount, objectively number one, Pareto dominant - rest on list-price arithmetic, while the same cluster documents that caching, turn count, retries and tool overhead dominate agentic bills and shows a case where a cheaper model finished within a cent of a pricier one. Grok 4.6 also trails on DeepSWE, Terminal-Bench and CursorBench, and its leaderboard rank is reported two different ways. The overstatement is moderate rather than severe because pricing, context window, availability and the turn-efficiency gap are genuinely documented.
Vendor-reported numbers and an acquirer vouching for the model
Much of the favourable material is interested: xAI reports its own index score, Cursor supplies the training-recipe and quality characterisations while its own CEO publicly endorsed the model, and SpaceX's announced but unclosed $60B all-stock acquisition of Cursor makes that endorsement a related-party statement. Nvidia's cost-advantage figures for its own router are Nvidia's own numbers, Box's CEO argues for value accruing to the routing layer his company sits in, and the discount arithmetic is from an investor. Two publishers explicitly label the vendor-reported material as such, which is why this is not a total blind spot.
Solid on the release, softer on the market thesis
Four publishers within about three days converge on the model's price, score and availability, and one supplies primary-documentation detail, so the factual core is reliable. Confidence is held back by the unresolved leaderboard conflict, several single-sourced numbers, the absence of any usage or spend data, and the fact that the cluster's most consequential claims - open weights capping closed-lab pricing, routing as the value layer - are interpretation rather than measurement.
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
build
Four frontier models in four days, and the cheapest number in your agent plan has an expiry date1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 12, 2026
1 article · August 13, 2026
1 article · August 12, 2026
1 article · August 14, 2026