Skip to content

Build6 publishers3 min readPublished

Google publishes Gemini 4 Argon's token prices before most teams can call the model

Google priced Gemini 4 Argon at $2 and $10 per million input and output tokens, then released it first to trusted cyber defenders in its Fairwind Program. Teams can budget against those rates now but cannot yet measure the token counts they multiply.

The Engineer · Build desk

Illustration accompanying Google publishes Gemini 4 Argon's token prices before most teams can call the model

What happened

  • The New Stack found Argon last of four models on FrontierSWE v2 and Terminal-Bench 4.0, despite Google's state-of-the-art 77.9% on DeepSWE v1.1.
  • Google says Fairwind participants and its internal teams are running the model without cyber guardrails.
  • After Fairwind, API customers and AI Ultra subscribers get access first, according to The New Stack, ahead of developers, enterprises and consumers.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Prompt layout becomes a pricing decision: a reused prefix bills at one-twentieth of the uncached rate, and teams outside Fairwind cannot yet measure how often theirs would hit the cache.
  • constraint Security teams cannot treat the 68% CWE-bench tie as a forecast for their own tooling, because each lab's score came from its own agent harness.
  • exposure Roadmaps built on an early general release carry schedule risk, since Google's only timing is 'as soon as possible' and this model already missed a June target, per The New Stack.

A cached input token on Argon costs 5% of the $2 rate, or $0.10 per million [3][1]. An uncached one costs 20 times as much [2][5]. An agent loop that resends the same long prefix on every turn pays somewhere between those two figures, and the cache hit rate decides where. Today only Fairwind participants and Google's own teams can measure one [1][18]. The pricing paragraph labels the rates introductory [2]. It does not say how long they last or what triggers the cache discount, and Google's only timing for wider access is "as soon as possible" [15].

Output is the other multiplier. Previous Gemini models stopped at 64,000 output tokens. Argon can generate up to one million [4]. At $10 per million, a response that fills the new ceiling costs $10 in output tokens alone [2]. At Argon's rate, a response held to the old 64,000 cap would have topped out at $0.64 [3]. Google wrote that "When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go." [5] A per-call spend limit sized for the old cap would need to rise about sixteenfold before one of those trajectories could finish [4].

Across Google's 18 comparison tests, The New Stack counts Argon first in 13 [19]. Google's coding headline is 77.9% on DeepSWE v1.1, which it calls a new state of the art [6]. The New Stack found Argon last of the four models on FrontierSWE v2 and Terminal-Bench 4.0 [7]. GPT-6 Astra and Opus 5.5 lead it there by 10.5 and nine points, respectively [7].

The security score depends on the harness. Argon ties GPT-6 Astra and Grok 4.7 at 68% on CWE-bench v1 [8]. The New Stack notes that the three major labs' models ran in their own agent harnesses, so that leaderboard scores each model and its tooling together [9]. For the 68% to predict anything in a team's own vulnerability pipeline, that pipeline would need to resemble the harness the score came from.

Knowledge work is where the lead holds up. Argon scores 51.3% on Zapier's AutomationBench, almost nine points ahead of Opus 5.5 [10]. On Harvey's Legal Agent Benchmark it scores 19.6%, nearly triple Fable 5.1. By The New Stack's count, that still means about one fully completed task in five [11].

The best engineering in the announcement is the libgav1 work. Argon agents took an existing Rust port of Google's open source video decoder and replaced 32K lines of SIMD code [12]. They ran rounds of profile-guided experiments and studied the compiler's output until safe Rust vectorized automatically [12]. The decoder now runs 2.7x faster than the Rust port, with identical video output [12]. A patient human performance engineer would do the same thing, more slowly. The work also happened on Google's code under Google's review process, where rewrites at this scale go through automated and manual auditing, emulation testing and review before production [13].

The model under test now differs from the one that will ship. Google says Fairwind participants and its internal teams get Argon without cyber guardrails [18]. It plans to iterate on guardrails with early testers before a wider release [15]. It is also taking part in the U.S. government's voluntary pre-release access process [16]. According to The New Stack, API customers and AI Ultra subscribers come next, followed by developers, enterprises and consumers [14]. The same outlet reports that Google first announced this model at I/O in May with a June launch planned, and shipped a series of Flash models instead [17].

What to watch

  • The terms behind the footnote on Argon's 'introductory' price, and whether the $2 and $10 rates change at general availability.
  • A date for API customers and AI Ultra subscribers, the first group The New Stack says gets access after Fairwind.
  • Independent coding results in third-party harnesses, especially on FrontierSWE v2 and Terminal-Bench 4.0, where Argon trails.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories