Invest2 distinct publishers3 min readUpdated
Google's budget tier now handles the summarize-and-compact work that fills agent invoices. The 75-cent introductory input rate lapses on December 31, 2026, and then input goes back to $1.50.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Google shipped Gemini 3.7 Flash on August 13, generally available in more than 160 countries on day one, at an introductory 75 cents per million input tokens and $3.75 per million output [1][2]. That is a promotional rate that holds until December 31, 2026, and Decrypt reports the input price doubles to $1.50 the day after [3][4], which puts the reset on January 1, 2027 [5] and means any 2027 inference budget sized off a current Flash invoice understates the run rate by a factor of two [6].
What makes that arithmetic matter is which jobs Flash does. Decrypt's description is the honest one: Flash is not the model you reach for when a problem is hard, it is what you use to sort text, compact agent sessions before they collapse under their own context, and summarize documents you do not want to pay a flagship to read [7]. Those are the line items that grow with agent count and turn count rather than with headcount, and they run against a one-million-token input window with 64,000 tokens out, across images, video, audio and PDFs, with tool calls and computer use [8].
On that narrow axis the upgrade is real. Decrypt's zero-shot test produced a playable browser game in 2 minutes 13 seconds, running correctly on the first attempt with collision and scoring logic intact [9]. Gemini 3.6 Flash, released July 21, could not produce a working file at all, and follow-up prompts asking it to repair its own malformed HTML went nowhere [10]; the reviewers handed the wreckage to DeepSeek, which found 11 bugs and shipped 8 fixes [11]. Google's own figure for coding efficiency moves to 43.6% from 34.4%, roughly a 27% relative gain inside a month [12]. The broader benchmark sheet, showing 3.7 Flash ahead of Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 categories, 1,588 Elo on Code Arena web development and 30.4% on AutomationBench, comes from Google's methodology and should be read as the company's claim [13].
The ceiling is where you would expect. Google positions the model explicitly as a fast coding tool rather than a reasoning model [14]. It failed Decrypt's bridge logic puzzle with the same wrong answer as Claude Fable 5 and set up a maths problem correctly without calculating it [15]; in the creative test it broke the one structural rule that decided the test, grasping the causal loop while still standing in the past [16]. Decrypt's summary is that outside the cheap-and-fast lane it gets outwritten by software you can download for free [17], specifically a community fine-tune of Qwen3.5-27B that runs on a single consumer GPU at no per-query cost [18].
The pricing detail worth internalising is that $1.50 per million input is not a new price. It is what 3.6 Flash cost, since 75 cents was described as half the predecessor's rate [19]. January is the end of a discount, not an increase, and it leaves you paying last generation's price for this generation's model. The 50% cut applied to output too, implying a $7.50 predecessor output rate [20], though Decrypt documents only the input reversion. Per job the absolute numbers stay small: a million tokens is about 750,000 words, roughly ten novels of code for under a dollar on the input side today [21], call it under two after the reset [22].
Watch three things. Whether Google extends the promotional rate past December 31, 2026, or lets it lapse quietly. Whether output pricing snaps back in step with input. And whether the three-week gap between 3.6 and 3.7, which CryptoBriefing reads as either fast internal iteration or a 3.6 launch that should have been held [23], means the model you priced in your budget is superseded before the discount ends.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Google shipped Gemini 3.7 Flash on August 13, generally available in more than 160 countries on day one.
The introductory rate for Gemini 3.7 Flash is $0.75 per million input tokens and $3.75 per million output tokens, a 50% cut from the previous release's pricing.
The promotional rate holds until December 31, 2026.
The model runs at 75 cents per million input tokens through December 31, half of 3.6 Flash's rate, before doubling to $1.50 on January 1.
Flash has never been the model for hard problems; it is used to sort text, compact agent sessions before they collapse under their own context, and summarize documents you do not want to pay a flagship to read.
Gemini 3.7 Flash takes up to a million input tokens, returns 64,000, reads images, video, audio and PDFs, and can call tools and drive a computer.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One independent hands-on review plus documented pricing, against vendor-supplied benchmark leads
Pricing terms, dates, availability and capability limits are stated consistently by both publishers, and Decrypt contributes reproducible first-party test results with named comparators. The quality leads, however, come from Google's own benchmark sheet — Decrypt flags this explicitly — and Cryptobriefing's coverage is aggregation of 9to5google and Decrypt rather than independent verification, so the strongest quality claims rest on a single vendor's methodology.
Broad day-one distribution, zero disclosed usage
Distribution is real and wide: day-one general availability in 160+ countries and exposure through the Gemini API, AI Studio, Antigravity and Gemini Spark. But no source reports a single customer deployment, token volume, revenue figure or migration off 3.6 Flash, and the only usage-adjacent data points are publisher test runs. Availability is measured; uptake is not observed.
Mildly overstated: vendor leads and game-generation framing outrun tested limits
The promotional layer — 11-of-18 category leads, 1,588 Elo, 30.4% AutomationBench, 43.6% coding efficiency, 'achieves playable game output' headlines — all originates with Google, while the independently tested picture shows a model that fails a logic puzzle, breaks a stated writing constraint, and loses to a free 27B fine-tune on a consumer GPU. The gap is modest rather than large because Decrypt discloses the vendor provenance, both publishers state plainly that it is not a reasoning model, and the price-expiry caveat is reported rather than buried.
Vendor-set benchmarks and a below-cost land-grab price with a built-in expiry
The commercial incentive is documented inside the cluster, not inferred: Cryptobriefing states that at $0.75 per million input tokens Google is not trying to recoup development costs but to establish Gemini as the default infrastructure layer, and that its provenance commitments could raise compliance costs for smaller rivals. The favourable quality figures come from Google's own methodology, and the discount lapses precisely once developers have built against it. On the publisher side, both Cryptobriefing pieces are aggregations of 9to5google and Decrypt.
Firm on price and dates, thin on verified capability and any real-world uptake
Confidence is anchored by facts both publishers state identically — release date, pricing pair, expiry, not-a-reasoning-model positioning — and by one publisher's reproducible tests. It is capped by effective single-publisher origination (Cryptobriefing is derivative), vendor-only benchmark evidence, and a complete absence of adoption or financial disclosure, leaving the post-January cost impact a projection rather than an observed outcome.
build
Gemini 3.7 Flash goes GA on one model layer, and that is the actual news1 distinct publisher
build
Four frontier models in four days, and the cheapest number in your agent plan has an expiry date1 distinct publisher
invest
Speed becomes a SKU: OpenAI and Google put a separate price on latency3 distinct publishers
product
A five-hour script beats Claude's watermark, so stop treating it as provenance3 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
cryptobriefing.com
2 articles · August 16, 2026
decrypt.co
1 article · August 16, 2026