Science1 distinct publisher2 min readUpdated
Artificial Analysis puts Grok 4.6 at 61 on its Intelligence Index, level with GPT-5.6 Sol, at $2/$6 per million tokens. The same pages record 48 seconds to first token.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
The arithmetic that matters most sits in the long-horizon results rather than the index score. Artificial Analysis measures Grok 4.6 resolving its Briefcase tasks in about 53 turns and 0.5B input tokens on average, against about 103 turns and 2.0B input tokens for Claude Opus 5 at max settings [8]. Four times fewer input tokens, bought at two-fifths the input price, is roughly a tenfold difference in input spend for work the same evaluator grades a tier lower [2]. Headline pricing cannot show you that, because token efficiency and per-token price multiply.
Output price is where the published gap is widest. Artificial Analysis argues that output tokens dominate cost in reasoning-heavy workloads [12], and $6 per million is a fifth of GPT-5.6 Sol's $30 and just under a quarter of Opus 5's $25 [1]. Measured against its own price class, the output rate is 40% below the $10 median [5], while the $2 input rate sits above the $1.75 median and is described on the evaluator's model page as somewhat expensive [14]. The discount is workload-shaped, and it favours jobs that read little and think a lot.
Two measurements bound the saving. Time to first token is 48.42 seconds, roughly seventeen times the 2.86-second median for comparable models [16][4], and output runs at 71 tokens per second against a 75 average [15]. The tidy blended figure of $1.35 per million assumes a 7:2:1 cache-hit-weighted mix [21], and the cache hit rate itself went up this generation, from $0.30 to $0.50 per million [9].
A caution on the evidence: both pages come from the same evaluator, and two of the strongest results, the AA-Briefcase Elo and the GDPval-AA v2 Elo, are its own benchmarks [7][4]. Its labelling is also unsteady. The identical figure of 72M output tokens generated during the index run is called somewhat verbose on one part of the page and better than average on another, against the same 72M median [18].
What the numbers support is narrower than a price war and more awkward for the vendors at the top. The gap from 61 to the 63 leader is smaller than the five points Grok 4.6 added over Grok 4.5 in a little over a month [3][2], so rank at this level is a perishable thing to sign a contract against. Holding headline pricing flat across a generation is unusual at the frontier, where gains have come with price rises [11]. That is the claim now needing a defence: not who scores highest, but what the extra $24 per million output tokens buys [1].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3.
Grok 4.6 headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, which Artificial Analysis describes as 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30).
Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5, with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. GDPval-AA v2 is Artificial Analysis's own measure of real-world agentic knowledge work.
Grok 4.6 cost $0.84 per Intelligence Index task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier.
Grok 4.6 debuts on AA-Briefcase, Artificial Analysis's private benchmark of long-horizon agentic knowledge work, with an Elo of 1577, at Fable 5-tier and behind the Claude Opus 5 family.
Grok 4.6 (high) has a time to first token of 48.42 seconds on SpaceXAI's API, at the higher end compared with a 2.86-second median for reasoning models in a similar price tier.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but single-evaluator, partly unauditable
The numbers are specific, dated and internally consistent on price and index score, and two of the agentic results (tau3-Banking, Terminal-Bench v2.1) sit on externally known benchmarks. Against that: every figure comes from one publisher that owns the scoring methodology, the two headline agentic rankings rest on its own GDPval-AA v2 and private AA-Briefcase with no published task set, the interpretive claims about frontier pricing norms and output-price dominance are asserted without data, and the model page contradicts itself on verbosity against an identical median. No independent replication exists in the cluster.
Available and benchmarked; no usage evidence
Adoption evidence stops at availability: the model is released on SpaceXAI's API with published pricing, and one evaluator has run it across its index suite. There is no deployment, customer, traffic, revenue or third-party integration disclosure anywhere in the cluster, so uptake is unmeasured rather than absent.
Cost-leadership framing outruns the measured record
The 61 index score, $2/$6 pricing and $0.84 per task are documented, so the core is not inflated. The overstatement is in framing: 'returns SpaceXAI to the intelligence frontier and leads on cost efficiency' is authored by the same party that owns the index, the two most flattering agentic results are private benchmarks, the 'pricing unchanged' narrative buries a 67% cache-hit price increase, and the publisher's own measurement of 48.42s to first token and below-median output speed is omitted from the analysis entirely.
Benchmark owner marketing its own indices and products
Artificial Analysis is both scorekeeper and commercial vendor here: the story is published on its own site, ranks the model on indices it owns, leans on two private benchmarks (GDPval-AA v2, AA-Briefcase) that only it can run, drives readers to its model page, and sits beside promotions for its Optima custom-benchmark platform and new Search Index. That is a clear interest in the salience of its rankings. There is no evidence in the cluster of payment by or a commercial relationship with SpaceXAI.
Concrete numbers, zero corroboration
Confidence is capped by structure, not vagueness: two sources, one publisher, one measurement pipeline, no independent evaluator or user report. Price, context window, index score and speed figures are consistent across the two pages, which supports the descriptive core; comparator rankings, private-benchmark Elos and the pricing-norm interpretation cannot be checked at all from what is supplied.
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
build
Four frontier models in four days, and the cheapest number in your agent plan has an expiry date1 distinct publisher
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
invest
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · August 22, 2026