Build1 publisher3 min readPublished
Grok 4.7 doubles cost per task at an unchanged per-token price
Grok 4.7 keeps Grok 4.6's $2/$6 token price yet costs $3.74 per task against $1.86, by Artificial Analysis' measurement. Teams that budget from the price sheet will undercount agent spend until they measure tokens per task on their own work.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- On Artificial Analysis' Intelligence Index, Grok 4.7 produced roughly 240 million tokens; Grok 4.6 needed 94 million, and the median model 88 million.
- Artificial Analysis clocked it at about 57 tokens per second, against 66 for Grok 4.6 and a median of 79, and called it "notably slow and very verbose".
- Its gains sit where xAI aimed: long office tasks and coding in xAI's own Grok Build harness, where it ranks 4th among coding agents.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Swapping Grok 4.6 for 4.7 on the same agent workload roughly doubles spend per task, and the return on that spend is two points on the Intelligence Index.
- decision The per-token column shows Grok 4.6 and 4.7 tied while their per-task cost differs twofold, so model selection has to run on measured cost per task at a fixed effort setting.
- constraint About 14% fewer tokens per second, spread over a longer answer, lengthens every task, so latency-bound uses such as interactive coding pay for the verbosity in wall-clock time as well as dollars.
A task's bill is the per-token price multiplied by the number of tokens the model chooses to write, and only the second term moved [1]. Divide Artificial Analysis' token counts and Grok 4.7 writes 2.55 times what Grok 4.6 did, and 2.73 times the median model [2].
The per-task cost multiple is 2.01 [1]. Output is only one of the two rates on the bill. If generated tokens rose 2.55 times and the total rose 2.01 times, spend on the $2 input side grew by less than 2.01 times [6]. A workload whose bill is mostly prompt will see a smaller multiple than one whose bill is mostly reasoning.
Reasoning effort changes the ratio. With both models pinned to xhigh, the output multiple is 2.13 [3]. That is below the mixed-effort figure and still above double. GPT-6 Astra, in the same Artificial Analysis article, writes about 27,000 output tokens per task, roughly a third of Grok 4.7's count [5].
xAI built the model to do this. Its launch post says the model "works longer on difficult tasks, checks its own work more carefully" [10]. The post also credits a "longer reinforcement learning run ... weighted toward problems that take many hours" [11]. Each Intelligence Index task took about 7.1 minutes [6].
The same post opens with "Twice as fast, at half the price of comparable models" [8]. A few lines later it says the model is "Served at the same price and speed as Grok 4.6" [9]. According to the dev.to write-up, the two sentences can both be true only if "twice as fast" is measured against another lab's model [16]. The price half holds per token against the flagships in xAI's own pricing table: $2 input is half of GPT-5.6 Sol Max's $4, and $6 output is 30% of its $20 [15][5]. Fable 5.1 Max lists at $10/$50 [15]. Per task, that comparison depends on how many tokens those models write on a given job.
Artificial Analysis' multiple is a measurement of its own task mix. It carries to another workload only if the tasks are long and multi-step like the index's, the effort setting matches, and the harness lets the model run until it decides it is done. The last condition is the hardest to predict. Developers on Hacker News, as quoted in the dev.to write-up, describe the model as lazy ("claim 'Done!'") and loopy ("Fix 80, Fix 81") at once [14]. An early "Done!" shrinks the token count, and a fix loop running into the eighties inflates it at $6 per million. I'd start a budget from the 2.13 equal-effort multiple, because it is the one figure here that holds the effort setting constant [3].
What to watch
- An Artificial Analysis cost-per-task figure for Grok 4.6 and 4.7 at matched effort, to replace the mixed-effort $3.74 against $1.86 comparison.
- Whether xAI documents which reasoning effort its API applies by default, since xhigh versus high changes the token multiple.
- Independent token counts on non-index workloads such as Grok Build coding runs, to test whether the 2.1x to 2.5x multiple holds outside the benchmark.