BuildWidely confirmed8 publishers3 min readPublished Updated
Cursor's task-cost table shrinks Grok 4.7's 7.5x list-price gap to about 1.9x
SpaceXAI lists Grok 4.7 at $2 and $6 per million tokens against $10 and $50 for GPT-6 Astra and Claude Fable 5.1, and the only per-task figures in the record put a Cursor run at $4.69 against about $9.
The Engineer · Build desk

What happened
- xAI, now operating under SpaceXAI, released Grok 4.7 on September 21 as its strongest model yet for coding and knowledge work, 40 days after Grok 4.6.
- Cursor's September high-effort test scored Grok 4.7 at 43.9% on CursorBench 4.0 with an average task cost of $4.69, against $5.20 for Grok 4.6 at 40.4%.
- On Terminal-Bench 4.0 the published numbers split: SpaceXAI reports 38% for Grok 4.7, while Artificial Analysis reports 26% against roughly 60% for GPT-6 Astra.
Why it matters
- cost The last increment of capability now has a measured price in one harness: about $0.81 per CursorBench point per task, paid by whoever runs the agent loop.
- decision A team comparing models cannot decide on the rate card alone, because the same two vendors sit 7.5x apart on tokens and about 1.9x apart on completed Cursor tasks.
- constraint Anyone wanting the doubled output speed has to work inside Cursor or Grok Build, since the fast version is not on the public API.
- contradiction A 12-point spread on one named benchmark means the harness configuration is deciding what the evaluation reports.
At list price, a million input tokens plus a million output tokens costs $8 on standard Grok 4.7 and $60 on standard GPT-6 Astra or Claude Fable 5.1 [17]. That is 7.5 times the price for the same mix [24]. No long agent run looks like that mix, and the one place in the record where somebody metered completed work shows the difference: in Cursor's September tests, Grok 4.7 averaged $4.69 per task while Fable 5.1 and Opus 5 cost about $9 [23][25]. The 7.5x becomes about 1.9x [28]. Cursor's published figures do not break out tokens consumed per task.
Same test, high effort on both sides: Grok 4.7 scored 43.9% on CursorBench 4.0, Fable 5.1 49.2%, Opus 5 44.7% [23][25]. Buying Fable's extra 5.3 points cost $4.31 more per task, about $0.81 per point [29]. Cursor cautions that small score differences may fall within normal evaluation variance [7]. Part of the increment a team is pricing may therefore be noise in the harness.
For $4.69 to transfer, a workload would have to resemble CursorBench 4.0 tasks, run at high effort, inside Cursor's harness, with Cursor's tool set and token budget. Change any of those and the scores move. SpaceXAI's own Terminal-Bench 4.0 setup puts Grok 4.7 at 38%, up from Grok 4.6's 20.3% [4]. Artificial Analysis puts it at 26% on the same named benchmark, against roughly 60% for GPT-6 Astra and 55% for Fable 5.1, with the cheaper DeepSeek V4.1 Flash at 27% [5][10]. theneuron.ai attributes divergence like that to differences in harnesses, reasoning settings, tool configurations, token budgets and evaluation environments [8].
The rate card has more variables than the two headline prices. SpaceXAI's developer documentation lists a 500,000-token context window, a May 2026 knowledge cutoff, and four reasoning-effort settings: low, medium, high and xhigh [9]. Effort is what moves token consumption, and SpaceXAI's own launch table compared 46.3% at xhigh against Grok 4.6's 40.4% at high [6]. The settings alone can explain that spread. The US regional endpoint adds a 10% token-price premium [11]. The fast version costs twice the standard rates for twice the output speed and runs only in Cursor and Grok Build [13].
Caching narrows the gap from the other side. Anthropic cut Fable 5.1's cache-read pricing to $0.25 per million tokens and says caching can reduce typical workload costs by roughly 25% and highly agentic workloads by as much as 45% [20]. Take the best case: 45% off $60 leaves $33, still just over four times Grok's $8 [26]. OpenAI also offers cached-input pricing [21].
On the independent Artificial Analysis Intelligence Index v4.3.2, which combines ten benchmarks, Grok 4.7 scores 46 while GPT-6 and Fable 5.1 lead at 53 each [14]. SpaceXAI reported 71% on DeepSWE v1.1 at high effort, close to GPT-5.6 Sol's 72.7% and above the 70% listed for Fable 5.1 [19]. SpaceXAI also says Grok 4.7 uses an entirely new safeguard stack, and reported 62.4% on LatchBio's biosafety benchmark and 3.3% of risky dual-use prompts getting through on HackerBench v0.3, its own cyber-safety test [18].
What to watch
- An independent cost-per-task table for Terminal-Bench 4.0 that publishes its harness and reasoning-effort settings would settle the 26 versus 38 percent split.
- Whether SpaceXAI publishes cache-read pricing for Grok 4.7, since the caching discounts in the record are Anthropic's and OpenAI's.
- Whether the fast version reaches the public API, and how Google's Gemini 3.8 Flash prices its long-horizon coding pitch against $2 and $6.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence60
- Adoption30
- Hype gap+25
- Incentives65
- Confidence58
Perspective Coverage
8 publishers- Builder
- Builder 51%
- Operator
- Operator 32%
- Investor
- Investor 17%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Standard API pricing for Grok 4.7 remains $2 per million input tokens and $6 per million output tokens, matching Grok 4.6.
- [2]
xAI, now operating under SpaceXAI, released Grok 4.7 on September 21 as its strongest model yet for coding and knowledge work.
- [3]
SpaceXAI said Grok 4.7 uses a larger base model than Grok 4.6 and received a longer reinforcement-learning run weighted toward problems that take hours to finish.
- [4]
On xAI's own Terminal-Bench 4.0 setup, Grok 4.7 jumped from Grok 4.6's 20.3% to 38%.
- [5]
The Decoder cites Artificial Analysis results putting Grok 4.7 at 26% on Terminal-Bench 4.0, compared with roughly 60% for GPT-6 Astra and 55% for Fable 5.1.
- [6]
SpaceXAI's launch table gives Grok 4.7 an xhigh score of 46.3% on CursorBench 4.0, up from Grok 4.6's 40.4% at high effort; those settings differ, limiting the usefulness of the direct comparison.
- [7]
Cursor cautions that small score differences may fall within normal evaluation variance.
- [8]
Different harnesses, reasoning settings, tool configurations, token budgets and evaluation environments can move agent benchmark scores around considerably, according to theneuron.ai.
- [9]
SpaceXAI's developer documentation lists a 500,000-token context window and a May 2026 knowledge cutoff for Grok 4.7, and developers can select low, medium, high or xhigh reasoning effort.
- [10]
On Terminal-Bench 4.0 the cheaper DeepSeek V4.1 Flash edges past Grok 4.7 at 27 percent.
- [11]
SpaceXAI offers a US regional API endpoint carrying a 10% token-price premium.
- [12]
Grok 4.7 is available through the SpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare, and becomes the default model in Grok Build.
- [13]
SpaceXAI offers a fast version of Grok 4.7 at twice the standard token rates with twice the output speed, limited to Cursor and Grok Build rather than the public API.
- [14]
On the independent Artificial Analysis Intelligence Index v4.3.2, which combines ten benchmarks, Grok 4.7 scores 46 and lands mid-pack while Claude Fable 5.1 and GPT-6 lead with 53 each.
- [15]
Both GPT-6 Astra and Claude Fable 5.1 list standard API pricing of $10 per million input tokens and $50 per million output tokens.
- [16]
In xAI's reported GDPval results, which are meant to approximate professional knowledge work, Grok landed between Fable 5.1 and GPT-6 Astra.
- [17]
At list price, one million input tokens plus one million output tokens costs $8 with standard Grok 4.7 and $60 with standard GPT-6 Astra or Fable 5.1.
- [18]
SpaceXAI says Grok 4.7 uses an entirely new safeguard stack, scored 62.4% on LatchBio's biosafety benchmark, and allowed 3.3% of risky dual-use prompts through on HackerBench v0.3, SpaceXAI's own cyber-safety test.
- [19]
Grok 4.7 scored 71% on DeepSWE v1.1 at high effort, close to GPT-5.6 Sol's 72.7% and above the 70% result listed for Fable 5.1.
- [20]
Anthropic cut Fable 5.1's cache-read pricing to $0.25 per million tokens and says caching can reduce typical workload costs by roughly 25% and highly agentic workloads by as much as 45%.
- [21]
OpenAI also offers cached-input pricing.
- [23]
Cursor's own September high-effort results scored Grok 4.7 at 43.9% against 40.4% for Grok 4.6, with an average task cost of $4.69 for Grok 4.7 versus $5.20 for Grok 4.6.
- [24]
The list-price ratio for that one-million-in, one-million-out mix is 7.5 times.
- [25]
In Cursor's test, Fable 5.1 and Opus 5 scored higher, at 49.2% and 44.7%, but cost about $9 per task.
- [26]
A 45% reduction on the $60 list cost of the one-million-in, one-million-out mix leaves $33, just over four times Grok 4.7's $8.
- [27]
Google's Gemini 3.8 Flash, released earlier this month, targets long-horizon software engineering.
- [28]
Measured on Cursor's average task cost, Fable 5.1 and Opus 5 cost about 1.9 times Grok 4.7 per task.
- [29]
In Cursor's test the extra 5.3 CursorBench points from Fable 5.1 cost $4.31 more per task, about $0.81 per point.
Sources
8 independent publishers whose own reporting we read for this story.
- blog.vercel.comGrok 4.7 now available and 40% off on AI Gateway, fx, and eve
1 article · September 20, 2026
- dev.toGrok 4.7 Is Not Chasing the Benchmark Crown—It Is Chasing Your Default Agent Slot
1 article · September 22, 2026
- runtimewire.comSpaceXAI ships Grok 4.7 at Grok 4.6 prices for longer agent work
1 article · September 21, 2026
- testingcatalog.comSpaceXAI releases Grok 4.7 for coding and knowledge work
2 articles · September 21, 2026
- the-decoder.comxAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
2 articles · September 21, 2026
- theneuron.aixAI’s Grok 4.7 Makes Frontier AI a Price War
2 articles · September 21, 2026
- thenewstack.ioGrok 4.7 was built to work for hours. It still fails most of the time.
2 articles · September 21, 2026
- SpaceXAI releases Grok 4.7, which it says is better at verifying its own work and managing longer context, available for $2/1M input and $6/1M output tokens
x.ai
1 article · September 21, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.