Skip to content

Product2 publishers3 min readPublished

Gemini 4 Argon raises Google's output cap from 64K to 1M tokens for single-pass agent jobs

Google's Gemini 4 Argon lifts its output limit to 1M tokens from 64K, aimed at agent jobs that finish in one pass. The higher ceiling raises both the most a single run can bill and the amount of output someone has to review before it ships.

The Product Desk · Product desk

Illustration accompanying Gemini 4 Argon raises Google's output cap from 64K to 1M tokens for single-pass agent jobs

What happened

  • Google plans to give trusted cyber defenders and its own internal teams a version of Argon that ships without cyber guardrails.
  • Wider release is described only as coming soon, starting with Google AI Ultra subscribers and paid API customers.
  • 9to5Google reports a launch price of $2 per million input tokens and $10 per million output, rising to $4 and $20 after an introductory period.
  • Inside Google, Argon agents applied memory optimizations across data centers that freed more than 300 TiB, with 500 TiB to 1 PiB expected in total.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Security teams have to establish which tier they would get before a benchmark tells them much, because Google reserves the full defensive capability for the guardrail-free version.
  • constraint A misalignment monitor that can stop execution means a long single-pass job can be halted partway, so pipelines built around one-go runs need checkpoints to resume from.
  • contradiction Ars Technica reported that Google had not announced API pricing while 9to5Google printed a full schedule, so a budget built on the $10 introductory output rate rests on one outlet's report.

Inside Google, Argon agents are rewriting the Fuchsia OS Zircon kernel from C/C++ into Rust. The kernel runs past 800,000 lines, and the rewrite goes through automated and manual auditing, emulation testing and review before any of it reaches production [8].

According to 9to5Google, the new headroom lets Argon generate hundreds of thousands of tokens in a single trajectory and solve tough problems in one go [2]. Ars Technica reports Google's version as completing more daunting tasks in a single step [15]. The new ceiling is 15.6 times the old one [1]. So far the users are thousands of Googlers, by Google's count [14], and the largest job Google describes goes to auditors before production [8]. In Google's own migration process, the pass is where review starts [8].

The output limit also caps what one call can bill. At the list price 9to5Google reported, a response that uses all 1M tokens costs $20 in output, against $1.28 for one that stops at the old 64K limit at the same rate [3][4]. At the introductory rate the full response costs $10 [2]. Failed runs bill as well. Google says Argon ranks first on Zapier's AutomationBench, which measures end-to-end execution across business functions, with a score of 51.3% [10]. An agent that goes wrong 900,000 tokens into a list-price run has spent $18 before anyone reads a line [5].

The cyber tiers sort buyers by who they are. The defenders using Argon today came in through Google's Fairwind Program [4]. On CWE-bench v1, which tests whether a model can remediate security vulnerabilities, Google says Argon ties for first at 68% [11]. Buyers outside the trusted tier get a model Google watches from the inside, with monitoring of internal activations aimed at cyber and CBRN misuse [12]. Google's announcement, as reported, does not say how a security team qualifies for Fairwind or what terms come with guardrail-free access.

In my view Argon is worth piloting first on a job that already produces more than 64K tokens of output and already has a human review step. Those teams gain the most from fewer chunks and are already paying for the reviewer. The tradeoff is that each failed run can bill up to $20 of output at list price before the reviewer sees it [3].

Two axes sort the decision. One is whether a job needs more than 64K tokens of output in a single response. The other is whether it is defensive security work a cyber guardrail could block. Short output outside security: little changes until a price holds past the introductory period. Long output outside security is the case Google built Argon for [1], and the budget line is per run, with $20 as the ceiling. Short security work turns on the tier, so the first conversation with Google is about Fairwind. Long security work needs a written price and Fairwind terms before any pilot result means much.

What to watch

  • Google publishing Argon's API price on its own pricing page, including how long the introductory period lasts.
  • Published criteria and terms for Fairwind Program access to the guardrail-free version.
  • Whether the Zircon kernel Rust rewrite clears Google's audits and reaches production.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories