Skip to content

Invest2 publishers3 min readPublished

xAI holds Grok's $2 token price for a model 40 Elo points behind Fable 5.1

Grok 4.7 arrived on Monday after five walked-back timelines, with 40% more parameters and Grok 4.6's list price intact. Cursor's own cost chart still puts its price per task above GPT-6 Astra and Claude Sonnet 5.

The Investor · Invest desk

Photograph accompanying xAI holds Grok's $2 token price for a model 40 Elo points behind Fable 5.1
Photo: techcrunch.com

What happened

  • xAI shipped Grok 4.7 on Monday after Musk walked the timeline back at least five times since late July, from "four weeks out" to "needs a few more days to cook" on September 11.
  • The model runs on 2.1 trillion parameters, up 40% from Grok 4.6's 1.5 trillion, with supplemental training on Starlink telemetry, SpaceX manufacturing records and engineering failure logs.
  • xAI held the starting price at $2 per million input tokens and $6 per million output, and sells a variant with twice the output speed at twice the price.
  • On the benchmarks xAI published, Grok 4.7 came second to Claude Fable 5.1 on GDPval and AA-Briefcase, and second to GPT-6 Astra on the EEBench electrical-engineering test.
  • Against its predecessor the model gained ground, scoring 46.3% on CursorBench 4.0 versus 40.4%, and 38% on Terminal Bench 4.0 versus 20.3%.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • decision Teams comparing list prices per token and teams comparing the cost of a finished task will pick different vendors from the same published numbers, so the choice of accounting unit now decides the contract.
  • cost By holding price while the base model grew by 600 billion parameters, xAI pays for the extra serving capacity itself instead of passing it to customers.
  • constraint With no dates attached to Grok 4.8, 4.9 or 5, anyone standardising on xAI has to justify the decision on the model that shipped Monday rather than the roadmap behind it.

Forty Elo points is the whole distance between Grok 4.7's 1695 on GDPval and Claude Fable 5.1's 1735 [4][1]. GDPval scores head-to-head on the system chess uses [27], so that gap converts to about 56 wins in 100 for Fable [2]. On AA-Briefcase the margin is 21 points, or roughly 53 in 100 [5][3].

The base model grew by 600 billion parameters. The list price did not move [8][4]. Musk called Grok 4.7 "a strong combination of intelligence, speed & low cost" on X [17], and xAI's own release note described it as "a notable improvement over Grok 4.6 at the same price and speed" [16].

Price is charged per token, and tokens are what the new model consumes more of. xAI says Grok 4.7 spends longer working through hard problems and double-checks its own answers more often than Grok 4.6 did [10]. On CursorBench 4.0, which plots accuracy against the cost and token count each task burns, Decrypt reports Grok 4.7 landing pricier per task than GPT-6 Astra and Claude Sonnet 5, with Fable 5.1 winning at every price point on the chart [11].

A team billed by consumption pays for the extra verification steps. The per-token discount can still leave the customer paying more per task.

The two published accounts of the same benchmark set do not agree. Crypto Briefing reports that xAI's results put Grok 4.7 ahead of GPT 5.6 Sol on several evaluations, with Sol and Fable 5.1 ahead on others [15]. Against Grok 4.6, the Terminal Bench 4.0 gain works out at 87% in relative terms, while CursorBench 4.0 moved 5.9 points [5][6]. Grok 4.7 also posted 71% on DeepSWE 1.1 at high reasoning effort [14], and, on the safety side, allowed 3.3% of risky dual-use prompts through on HackerBench 0.3 and scored 62.4% on a LatchBio biosafety benchmark [23].

The months of delay went into a larger base model and a longer reinforcement learning run on tasks that take hours to finish [22]. They did not go into multimodal, which Musk said still needs work when he set expectations at "roughly on par with" Anthropic's Claude Opus 5.0, not the newer Opus 5.1 [18]. He then sketched Grok 4.8 as a meaningful step up, Grok 4.9 in "Astra/Fable class" and Grok 5 as a possible frontier leader, with no release dates. "We shall see," he wrote [19].

I'd expect distribution to decide more of this than either leaderboard. Grok 4.7 shipped with no waitlist into the Grok app, Cursor, Grok Build and the xAI API [20], and reaches developers through third-party coding tools, model routers and cloud platforms [21]. The counter-case is the one Cursor's chart makes: agent work bills by consumed tokens, and a model that reasons longer gives back its token-price advantage exactly where the spend concentrates. An independent cost-to-completion measurement at matched accuracy would settle it. If Grok 4.7 finishes the same task for less money, Decrypt's claim that xAI has consistently undercut Anthropic and OpenAI on cost per token holds [25]; if it does not, xAI is selling second place at a premium per task.

What to watch

  • An independent cost-per-completed-task measurement at matched accuracy, testing whether the $2 per million token price survives task-level accounting.
  • A dated release for Grok 4.8, or for the Grok 4.9 that Musk placed in "Astra/Fable class."
  • What the selected cybersecurity partners given red-team access report against the 3.3% HackerBench 0.3 pass-through figure.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories