Invest3 publishers3 min readPublished
Google grades its own Gemini 4 Argon 3.7 points ahead of Claude on long-horizon coding
Google graded its own Gemini 4 Argon at 77.9% on the DeepSWE coding test, 3.7 points above Claude Opus 5.5's leaderboard score. When paying API customers get the model, they will pay half Google's list price for a period the company has not set.
The Investor · Invest desk

What happened
- Google's Fairwind Program, launched September 2 with more than 650 partners including governments, gets Argon first and without its usual cyber guardrails.
- In Google's own comparison table, Argon leads on 12 of 18 benchmarks, ties one and trails on five.
- Argon arrived one week after Anthropic's Claude Opus 5.5 and one day after OpenAI's GPT 6.1 Sol.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- cost Buyers who build on the introductory rate face a doubling to $4 and $20 per million tokens on a date Google has not set, taking a maximum-length reply from $10 to $20 in output.
- exposure Exploit-writing ability 12.7 points above Google's restricted cyber model reaches more than 650 outside organisations before Google has scaled the safeguards it wants for a public launch.
- decision Teams picking a coding model on DeepSWE must either trust a 3.7-point gap Google graded itself or wait until Argon is graded on the same public leaderboard as its rivals.
Google computed Argon's 77.9% on DeepSWE itself, according to Decrypt [1][3]. The 74.2% for Claude Opus 5.5 and the 74.1% for GPT-6 Astra came from a public leaderboard and company reports [2][3]. So the lead over Anthropic's model is 3.7 points [1], and two different graders produced the scores on either side of it. Google's own table also concedes ground: Argon leads 12 of 18 benchmarks, ties one and trails on five [5].
The cyber claim changes with the publisher. Decrypt's headline says Argon "tops all other AI models on cybersecurity" [8]. CNBC reports Google saying the model ties for first, level with GPT-6 Astra and Grok 4.7 on cybersecurity evaluation benchmarks [9]. Argon's clearest lead is on Gray Swan's indirect prompt injection test, where attacks succeeded 0.7% of the time within 15 tries, against 1.0% for both Claude Opus 5.5 and Claude Fable 5.1 [6]. The gap is 0.3 points, or 30% fewer successful attacks [3]. GPT-6 Astra was tricked 8.5% of the time [7].
Revenue comes last in the release order. More than 650 Fairwind partners, including governments and critical infrastructure operators, get the model first and "without cyber guardrails" [10]. Paid API customers and Google AI Ultra subscribers come after [12]. When the API opens, the introductory rate of $2 per million input tokens and $10 per million output tokens is half the standard $4 and $20, and Google has not said when the discount ends [13][4]. The output ceiling is now 1 million tokens per reply, up from 64,000, a 15.6-fold rise [14][6]. One maximum-length reply costs $10 in output at the introductory rate and $20 at list [5].
The saving Google can already count sits inside its own fleet. The company told CNBC that Argon has been used to optimise data-center memory, freeing hundreds of terabytes without buying more hardware [16]. Argon follows a July in which Google shipped smaller Flash models but skipped the promised Gemini 3.5 Pro; Alphabet shares fell about 4.4% [17].
"Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible," Tulsee Doshi, Google's Gemini model product lead, told CNBC [18]. On the Wiz Penetration Test Benchmark, an internal Google test of writing working exploits against web-application flaws, Argon solved 70.9% on the first try [15]. Google's restricted Gemini 3.8 Flash Cyber solved 58.2%, so the gain is 12.7 points [15][7]. Google says it will scale safeguards in four areas, including misuse and prompt injection, before launching publicly [19].
I think the coding lead points the right way but is small enough that a neutral grader could erase it. The case against that caution is the 28.9-point gain over the 49% Gemini 3.6 Flash scored in July [4][2]. A jump that size is hard to put down to grading method. If the public leaderboard confirms something near 77.9%, the caution is wrong and Google holds the top score until the next release. Claude Opus 5.5 shipped a week before Argon, and GPT 6.1 Sol a day before [11]. If the leaderboard puts Argon below 74.2%, the coding lead is gone and the gain over Flash is what remains.
What to watch
- The date Google ends the $2 and $10 introductory pricing and moves Argon to the $4 and $20 standard rates.
- When paid API customers and Google AI Ultra subscribers get access, since that depends on Google scaling safeguards in four areas including misuse and prompt injection.
- Whether Fairwind partners publish vulnerabilities found with Argon, as Mozilla did with the 271 Firefox flaws an early Claude Mythos helped find.