Skip to content

Invest3 publishers3 min readPublished

Google grades its own Gemini 4 Argon 3.7 points ahead of Claude on long-horizon coding

Google graded its own Gemini 4 Argon at 77.9% on the DeepSWE coding test, 3.7 points above Claude Opus 5.5's leaderboard score. When paying API customers get the model, they will pay half Google's list price for a period the company has not set.

The Investor · Invest desk

Illustration accompanying Google grades its own Gemini 4 Argon 3.7 points ahead of Claude on long-horizon coding

What happened

  • Google's Fairwind Program, launched September 2 with more than 650 partners including governments, gets Argon first and without its usual cyber guardrails.
  • In Google's own comparison table, Argon leads on 12 of 18 benchmarks, ties one and trails on five.
  • Argon arrived one week after Anthropic's Claude Opus 5.5 and one day after OpenAI's GPT 6.1 Sol.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • cost Buyers who build on the introductory rate face a doubling to $4 and $20 per million tokens on a date Google has not set, taking a maximum-length reply from $10 to $20 in output.
  • exposure Exploit-writing ability 12.7 points above Google's restricted cyber model reaches more than 650 outside organisations before Google has scaled the safeguards it wants for a public launch.
  • decision Teams picking a coding model on DeepSWE must either trust a 3.7-point gap Google graded itself or wait until Argon is graded on the same public leaderboard as its rivals.

Google computed Argon's 77.9% on DeepSWE itself, according to Decrypt [1][3]. The 74.2% for Claude Opus 5.5 and the 74.1% for GPT-6 Astra came from a public leaderboard and company reports [2][3]. So the lead over Anthropic's model is 3.7 points [1], and two different graders produced the scores on either side of it. Google's own table also concedes ground: Argon leads 12 of 18 benchmarks, ties one and trails on five [5].

The cyber claim changes with the publisher. Decrypt's headline says Argon "tops all other AI models on cybersecurity" [8]. CNBC reports Google saying the model ties for first, level with GPT-6 Astra and Grok 4.7 on cybersecurity evaluation benchmarks [9]. Argon's clearest lead is on Gray Swan's indirect prompt injection test, where attacks succeeded 0.7% of the time within 15 tries, against 1.0% for both Claude Opus 5.5 and Claude Fable 5.1 [6]. The gap is 0.3 points, or 30% fewer successful attacks [3]. GPT-6 Astra was tricked 8.5% of the time [7].

Revenue comes last in the release order. More than 650 Fairwind partners, including governments and critical infrastructure operators, get the model first and "without cyber guardrails" [10]. Paid API customers and Google AI Ultra subscribers come after [12]. When the API opens, the introductory rate of $2 per million input tokens and $10 per million output tokens is half the standard $4 and $20, and Google has not said when the discount ends [13][4]. The output ceiling is now 1 million tokens per reply, up from 64,000, a 15.6-fold rise [14][6]. One maximum-length reply costs $10 in output at the introductory rate and $20 at list [5].

The saving Google can already count sits inside its own fleet. The company told CNBC that Argon has been used to optimise data-center memory, freeing hundreds of terabytes without buying more hardware [16]. Argon follows a July in which Google shipped smaller Flash models but skipped the promised Gemini 3.5 Pro; Alphabet shares fell about 4.4% [17].

"Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible," Tulsee Doshi, Google's Gemini model product lead, told CNBC [18]. On the Wiz Penetration Test Benchmark, an internal Google test of writing working exploits against web-application flaws, Argon solved 70.9% on the first try [15]. Google's restricted Gemini 3.8 Flash Cyber solved 58.2%, so the gain is 12.7 points [15][7]. Google says it will scale safeguards in four areas, including misuse and prompt injection, before launching publicly [19].

I think the coding lead points the right way but is small enough that a neutral grader could erase it. The case against that caution is the 28.9-point gain over the 49% Gemini 3.6 Flash scored in July [4][2]. A jump that size is hard to put down to grading method. If the public leaderboard confirms something near 77.9%, the caution is wrong and Google holds the top score until the next release. Claude Opus 5.5 shipped a week before Argon, and GPT 6.1 Sol a day before [11]. If the leaderboard puts Argon below 74.2%, the coding lead is gone and the gain over Flash is what remains.

What to watch

  • The date Google ends the $2 and $10 introductory pricing and moves Argon to the $4 and $20 standard rates.
  • When paid API customers and Google AI Ultra subscribers get access, since that depends on Google scaling safeguards in four areas including misuse and prompt injection.
  • Whether Fairwind partners publish vulnerabilities found with Argon, as Mozilla did with the 271 Firefox flaws an early Claude Mythos helped find.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories