Invest1 publisher2 min readPublished
Grok 4.7 buys 5.9 points of CursorBench at Grok 4.6's token price
xAI built Grok 4.7 on a larger base model with a longer reinforcement-learning run and kept the API at Grok 4.6's rates. The open question for buyers is how many tokens the longer runs burn.
The Investor · Invest desk

What happened
- On CursorBench 4.0, which covers longer software-engineering tasks, Grok 4.7 scored 46.3 percent against 40.4 percent for Grok 4.6.
- Standard API pricing is unchanged from Grok 4.6 at $2 per million input tokens and $6 per million output, with a faster serving tier sold at twice the output speed and twice the price.
- The model is live in Cursor and in Grok Build, and is also sold through the Grok API, third-party coding harnesses, model routers and cloud platforms.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- cost Any extra tokens Grok 4.7 spends staying on a task are billed to the buyer at the same output rate, so an unchanged rate card can still produce a larger invoice.
- decision Existing Grok users in editors and agent stacks can settle the upgrade by metering output tokens on one repeated task, since the per-token price is identical either way.
- contradiction The largest published gain sits on xAI's own harness while the launch account places the model behind other frontier systems on some evaluations, so a buyer's own reproduction can land lower.
- constraint If capability keeps arriving at a constant rate card, xAI's remaining competitive lever on the standard tier is the rate itself.
The score moved 5.9 points and the rate card did not move at all, so on the standard tier xAI now charges about 12.7% less per point of CursorBench 4.0 than it did for Grok 4.6 [1][2][12]. That is the seller's version. The buyer's version is cost per finished task, and it runs the other way if the model does what xAI describes, which is stay with demanding assignments for longer and review its own output more carefully [3]. Persistence is output tokens. At $6 per million, a run that emits 15% more tokens bills 15% more [12][8].
The capability came out of the training budget. Grok 4.7 sits on a larger base model than 4.6 and got a longer reinforcement-learning run, weighted toward problems that can take many hours to finish instead of short prompts [4][5]. xAI did not publish what that run cost. The company frames the release as a practical upgrade over 4.6 and not a change in pricing or serving speed [2].
The gains are uneven. Terminal-Bench 4.0 went from 20.3% to 38.0% on xAI's reported harness, an 87% relative move [8][3]; EEBench went from 53.0 to 64.0, or 20.8% [9][4]; CursorBench 14.6% [1]; DeepSWE v1.1 from 65.2% to 71.0% at high effort, 8.9% [7][7]; and AA Briefcase v1.1, the multi-hour office test, from 1,546 to 1,657, or 7.2% [10][5]. The account of the launch says those results put Grok 4.7 near other frontier systems on some tests and behind them on others, depending on the evaluation and the effort setting [11].
Context stays at 500,000 tokens, with text and image input, tool use, search and code execution [15]. Filling that window once costs a dollar at $2 per million input tokens [6]. The model was also trained to work natively with the Grok Bot harness, which xAI expects to help with conversation, document drafting and general office-style work [17].
Three ways this goes. If output per completed task holds flat, the unchanged rate card is a real price cut and the 12.7% per point is what buyers collect [2]. If the extra persistence pushes output per task up by more than the 14.6% score gain, buyers are paying more for better work [1]. And since 4.7 landed only weeks after 4.6 [18], the rate card has a decent chance of moving before the next benchmark table does. I'd expect the middle case, since self-checking and long-horizon persistence are both token-spending behaviours [3]. The check is one repository task run on both models, with the comparison made on output token counts instead of scores.
What to watch
- Third-party reproductions of the Terminal-Bench 4.0 move from 20.3% to 38.0% outside xAI's reported harness.
- Whether xAI publishes output tokens per completed task, the figure that decides if flat pricing is a cut.
- The gap to the next Grok release: 4.7 followed 4.6 by weeks, so the rate card may move before the scores do.