Skip to content

Build1 publisher2 min readPublished

Gemini 3.8 Flash beats partner-only Argon for anything shipping this quarter

Google prices Gemini 3.8 Flash at $0.75/$3.75 per million input/output tokens through 2026, three-eighths of what partner-only Gemini 4 Argon costs. Building on Flash now works if the later Argon swap moves the thinking settings along with the model name.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Gemini 3.8 Flash beats partner-only Argon for anything shipping this quarter
Generated illustration

What happened

  • Argon's introductory price of $2 per million input tokens and $10 per million output has no announced end date and rises to $4 and $20 afterward.
  • Flash caps a single response at 65,536 tokens, against the 1M-token output limit Google lists for Argon.
  • For a call with 100,000 input and 8,000 output tokens run 1,000 times a day, the post puts Flash at $105 a day and Argon at $280 at introductory rates.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Flash's launch rate ends with 2026, after which the example workload costs $210 a day, double what a launch-priced budget assumes.
  • constraint A 200,000-token answer has to be split across at least four Flash calls, so long-output features become multi-call jobs on Flash that Argon's stated limit would return in one response.
  • contradiction Vals AI's 262,000-token output ceiling for the Argon setup it tested is about a quarter of Google's 1M figure, so a design counting on near-1M single responses rests on Google's number alone.
  • exposure Prototyping on Flash's free tier with real customer prompts puts that data into Google's product improvement, by Google's own description of the tier.

A comparison posted on dev.to recommends shipping on Flash and keeping Argon one configuration switch away [24]. It builds that switch as a shell default, `MODEL="${GEMINI_MODEL:-gemini-3.8-flash}"`, interpolated into the `generateContent` URL [22]. It also warns against hardcoding a guessed Argon ID, advice that should not need printing [21]. The same request body pins `thinkingConfig.thinkingLevel` to `"medium"` and `maxOutputTokens` to 8192 [22].

I agree with building on Flash. I would also put the thinking level and the output cap in the same environment as the model name, so one deploy changes all three. Flash's levels are described in Google's thinking documentation [25]. Google has not published Argon's model ID, input window, thinking levels or default, how it bills thinking tokens, or a batch price [21][6][8][15][20]. Argon's evaluations ran at the highest thinking settings [8].

The post's daily cost comparison holds token counts equal across the two models [17]. On current Gemini models, Flash included, thinking tokens bill as output, and the example's 8,000 output tokens count them [15]. The post's Argon figures assume Argon bills the same way [15]. Google says higher Flash thinking levels can consume more tokens [18]. Suppose Argon runs at the setting its benchmarks used and emits more output tokens per call than the example assumes. Then the per-call gap grows past the 8/3 per-token ratio [16][8]. The post's test script already carries the right guard: it asserts `usageMetadata.thoughtsTokenCount` stays below 20,000 [23].

Argon's benchmark case rests on two rows. Both are Google-reported, Argon's methodology attributes both to Vals AI, and the two tables appeared a month apart against different comparison models [10]. On finance, Argon scores 65.4% to Flash's 61.4% [11]. For the legal-agent gap to carry over to someone else's agent, two things have to be true. The task has to resemble the benchmark's, and Argon has to run at the setting that produced the score [12][8]. Arena's text leaderboard, a third-party ranking, places Argon (High) first and Flash (High) eleventh on preliminary votes [13].

Flash has one more discount. Batch requests are 50% off, at $0.375 input and $1.875 output per million tokens through 2026, according to the Gemini API pricing page as cited in the post [20]. On batch, the example workload comes to $52.50 a day [3], for jobs that can run as batch work.

For anyone outside Fairwind, access settles it. Argon is open only to Fairwind Program partners [5]. Even Flash's cyber sibling, Gemini 3.8 Flash Cyber, is GA only behind an allowlist on the Gemini Enterprise Agent Platform [14]. Everyone else can build today on the general Gemini 3.8 Flash [1].

What to watch

  • Google publishing Argon's model ID, thinking levels and thinking-token billing, the inputs a working config switch needs.
  • An end date for Argon's $2/$10 introductory pricing, or access opening beyond the Fairwind Program.
  • Independent tests of how much output Argon actually returns in one response, measured against Vals AI's 262,000-token configuration.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories