Skip to content

Invest2 publishers3 min readPublished

Gemini 3.7 Flash Is Cheap Until January 1, When The Agent Bill Doubles

Google's budget tier now handles the summarize-and-compact work that fills agent invoices. The 75-cent introductory input rate lapses on December 31, 2026, and then input goes back to $1.50.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Gemini 3.7 Flash Is Cheap Until January 1, When The Agent Bill Doubles
Photo: decrypt.co

What happened

  • Google shipped Gemini 3.7 Flash on August 13, generally available in more than 160 countries on day one.
  • The introductory rate for Gemini 3.7 Flash is $0.75 per million input tokens and $3.75 per million output tokens, a 50% cut from the previous release's pricing.
  • The promotional rate holds until December 31, 2026.
  • The model runs at 75 cents per million input tokens through December 31, half of 3.6 Flash's rate, before doubling to $1.50 on January 1.
  • The input price reset to $1.50 per million tokens falls on January 1, 2027.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

Google shipped Gemini 3.7 Flash on August 13, generally available in more than 160 countries on day one, at an introductory 75 cents per million input tokens and $3.75 per million output [1][2]. That is a promotional rate that holds until December 31, 2026, and Decrypt reports the input price doubles to $1.50 the day after [3][4], which puts the reset on January 1, 2027 [5] and means any 2027 inference budget sized off a current Flash invoice understates the run rate by a factor of two [6].

What makes that arithmetic matter is which jobs Flash does. Decrypt's description is the honest one: Flash is not the model you reach for when a problem is hard, it is what you use to sort text, compact agent sessions before they collapse under their own context, and summarize documents you do not want to pay a flagship to read [7]. Those are the line items that grow with agent count and turn count rather than with headcount, and they run against a one-million-token input window with 64,000 tokens out, across images, video, audio and PDFs, with tool calls and computer use [8].

On that narrow axis the upgrade is real. Decrypt's zero-shot test produced a playable browser game in 2 minutes 13 seconds, running correctly on the first attempt with collision and scoring logic intact [9]. Gemini 3.6 Flash, released July 21, could not produce a working file at all, and follow-up prompts asking it to repair its own malformed HTML went nowhere [10]; the reviewers handed the wreckage to DeepSeek, which found 11 bugs and shipped 8 fixes [11]. Google's own figure for coding efficiency moves to 43.6% from 34.4%, roughly a 27% relative gain inside a month [12]. The broader benchmark sheet, showing 3.7 Flash ahead of Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 categories, 1,588 Elo on Code Arena web development and 30.4% on AutomationBench, comes from Google's methodology and should be read as the company's claim [13].

The ceiling is where you would expect. Google positions the model explicitly as a fast coding tool rather than a reasoning model [14]. It failed Decrypt's bridge logic puzzle with the same wrong answer as Claude Fable 5 and set up a maths problem correctly without calculating it [15]; in the creative test it broke the one structural rule that decided the test, grasping the causal loop while still standing in the past [16]. Decrypt's summary is that outside the cheap-and-fast lane it gets outwritten by software you can download for free [17], specifically a community fine-tune of Qwen3.5-27B that runs on a single consumer GPU at no per-query cost [18].

The pricing detail worth internalising is that $1.50 per million input is not a new price. It is what 3.6 Flash cost, since 75 cents was described as half the predecessor's rate [19]. January is the end of a discount, not an increase, and it leaves you paying last generation's price for this generation's model. The 50% cut applied to output too, implying a $7.50 predecessor output rate [20], though Decrypt documents only the input reversion. Per job the absolute numbers stay small: a million tokens is about 750,000 words, roughly ten novels of code for under a dollar on the input side today [21], call it under two after the reset [22].

Watch three things. Whether Google extends the promotional rate past December 31, 2026, or lets it lapse quietly. Whether output pricing snaps back in step with input. And whether the three-week gap between 3.6 and 3.7, which CryptoBriefing reads as either fast internal iteration or a 3.6 launch that should have been held [23], means the model you priced in your budget is superseded before the discount ends.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories