Skip to content

Product1 publisher2 min readPublished

SpaceX prices Grok 4.7 at $4.69 a task on a benchmark it owns

SpaceX says Grok 4.7 averages $4.69 per task on CursorBench 4.0, which was built by Cursor, one of its own acquisitions. The rivals it undercut were running quality-first configurations, and GPT-6 Astra beat it on chip design.

The Product Desk · Product desk

Photograph accompanying SpaceX prices Grok 4.7 at $4.69 a task on a benchmark it owns
Photo: yahoo.com

What happened

  • SpaceX says Grok 4.7 finished the challenges in CursorBench 4.0 at an average cost of $4.69 per task, which put it ahead of GPT-5.6 Sol and Fable 5.1.
  • CursorBench 4.0 was developed by Cursor, one of SpaceX's recent acquisitions, so the headline cost figure comes from a benchmark the model's owner also owns.
  • List pricing starts at $2 per million input tokens and $6 per million output tokens, with a version for latency-sensitive work that runs twice as fast for twice the price.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision For a team pinning a default coding model, $4.69 per task is the only number in this release that maps onto a monthly invoice, so reproducing it on internal work is what stands between the claim and a budget line.
  • constraint With no accuracy figure beside the cost, cost per completed task including retries cannot be worked out from what SpaceX published, and the cheapest model per attempt can still be the dearest per merged change.
  • contradiction The same evaluation has Grok 4.7 ahead of Fable 5.1 and behind GPT-6 Astra on EEBench, so a legal ops buyer and a chip design buyer cannot both read "most capable" the same way.
  • cost Teams that need the faster tier pay double for the same work: $9.38 per task at the CursorBench token mix.

A platform lead setting the default model in a shared Cursor config is picking a completion rate and a cost per task at the same time. SpaceX published one of the two. Its evaluation puts Grok 4.7's average cost on CursorBench 4.0 at $4.69 per task [3]. It did not publish an accuracy or completion score for that benchmark [12].

At list prices of $2 per million input tokens and $6 per million output tokens [7], $4.69 covers about 782,000 output tokens if a task were pure output, or about 2.35 million input tokens if it were pure input [1]. A real run mixes both. So the average CursorBench task is a long multi-step job somewhere above a million tokens.

That fits what the model is built to run inside. Grok 4.7 works with the Grok Bot harness, which splits complex work among several agents that run in parallel and verify each other's output [9]. Cross-checking adds tokens to the meter, and the $4.69 includes them.

CursorBench 4.0 was developed by Cursor, which SpaceX acquired [2]. The rival entries Grok 4.7 came in under, GPT-5.6 Sol and Fable 5.1, were hardware-intensive versions that prioritise output quality over cost-efficiency, according to SiliconANGLE's account of the evaluation [4]. I would want the cost-efficient configurations of both rivals in that table before I read the ranking as a price win.

The capability picture splits by task. Grok 4.7 outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and on EEBench, the chip design set [5], and its EEBench score fell behind OpenAI's GPT-6 Astra [6]. A legal ops team and a silicon team reading the same announcement get different answers.

There is also a second price tier. SpaceX offers a version that processes prompts twice as fast and costs twice as much [8], which works out to $4 and $12 per million tokens [2] and $9.38 per task at the same token mix [3]. Grok 4.7 arrived less than a week after Grok Voice Transcribe 2.0, which SpaceX says doubles its predecessor's accuracy at half the cost [11].

The test that settles a default is cost per accepted change on your own repository, measured over a week of real tickets, at whichever tier you would actually ship, with failed attempts counted in the numerator. The $4.69 came out of SpaceX's own run on a benchmark SpaceX owns [2][3].

What to watch

  • Whether SpaceX publishes an accuracy or completion rate for Grok 4.7 on CursorBench 4.0 next to the $4.69 cost figure.
  • Whether OpenAI or Fable's makers post CursorBench 4.0 cost-per-task numbers using cost-efficient configurations of their models.
  • Whether the doubled-price low-latency version becomes the default tier inside Cursor itself.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories