Science3 publishers3 min readPublished
Gemini 4 Argon costs 2.7 times as much per task as GPT-6.1 Sol at the same token price
Google's Gemini 4 Argon matches GPT-6.1 Sol's $2/$10 token price but costs 2.7 times as much per task, according to Artificial Analysis. Argon uses more tokens per job, so buyers still have to compare frontier models by cost per completed task.
The Scientist · Science desk

What happened
- The $2/$10 rate is a 50% launch discount on standard pricing of $4/$20, and Google has not confirmed when the promotion ends.
- On AA-Omniscience, Argon's hallucination rate is 15%, against 51% for GPT-6 Astra and 54% for GPT-6.1 Sol.
- Access is limited to more than 650 cyber defenders in Google's Fairwind Program, for defensive use only.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- cost Budgets set at the launch rate will understate spend, because Argon's per-task cost doubles whenever Google ends the promotion.
- decision Choosing between Argon and Sol has to come down to cost per completed task on the buyer's own workload, and on the public index Sol is far cheaper by that measure.
- capability Workloads where a confident wrong answer is expensive now have a frontier option that declines to guess far more often than OpenAI's current models do.
- constraint Most enterprises cannot yet test Argon in production, so the only independent per-task cost evidence available to them comes from a single benchmarking firm.
Token rates are one input to a bill. The other is how much the model writes, and Artificial Analysis found Argon wordy. Across the full Intelligence Index, Argon generated 110 million tokens, against a median of 82 million for reasoning models in its price tier [18]. Per task, it writes about 2.3 times as many output tokens as GPT-6 Astra [2].
If the two models really charge the same rates, the gap with Sol comes from how many tokens each one uses. Argon costs $1.99 a task, 2.7 times Sol's figure, which puts Sol at about 74 cents a task [4][1]. That fits the firm's earlier finding that GPT-6.1 Sol runs at less than a quarter of GPT-6 Astra's cost per task. A quarter of Astra's $3.26 is about 82 cents [19][5]. Against Astra, Argon comes in at 60% of the per-task cost [4].
Forkast reported that Sol launched at the same rates a day before Argon [2], and it called the match "the commodity floor arriving in the frontier model tier" [12]. That floor holds only while Google's promotion runs. When it ends, Artificial Analysis expects Argon to cost $3.98 a task [5]. If OpenAI keeps Sol's price where it is, Argon would then cost about 5.4 times as much per task as Sol: the 2.7 multiple, doubled [3].
The evidence on what the extra money buys is mixed. Argon's headline coding score, 77.9% on DeepSWE v1.1, is Google's own figure, and Forkast notes that Google has not published independent verification [13]. The independent index behind the cost figures combines reasoning, knowledge, mathematics and coding [21].
The hallucination result is the most interesting number in the release, and it comes with lower accuracy. Argon gets 50% of AA-Omniscience questions right, 13 points below GPT-6 Astra. Its overall score of 42 sits beside Astra's 43 and Sol's 42 [9]. So Argon answers fewer questions correctly, but it guesses wrong far less often [8][9]. The thing this doesn't tell you is whether your workload is better served by a model that declines to answer.
The output limit is now 1 million tokens, up from 64,000 [16]. That changes what a single call can cost. A Gemini API feature called Long Decode Continuation pauses long responses and resumes them across follow-up calls, so reasoning can run that long without hitting request timeouts [17]. One full million-token response costs $10 in output tokens at the launch rate and $20 at standard pricing [4]. Even the context window is in dispute: Forkast reports 2 million tokens [14], while Artificial Analysis lists 1 million [15].
On a wider release, Google wrote that it will "continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible" [20].
What to watch
- Google confirming an end date for the 50% promotion, which moves Argon's per-task cost from $1.99 to $3.98.
- Independent DeepSWE v1.1 results for Argon against Claude Opus 5.5 and GPT-6 Astra, replacing Google's self-reported 77.9%.
- Any change to Sol's rates, or to Argon's token use at general release, since either would change the 2.7x per-task gap.