Build4 distinct publishers3 min readPublished
Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Billing is per token, not per task. That is what makes the unchanged rate card misleading. Google's own copy says 3.8 Flash executes extra reasoning steps and calls tools iteratively, and may use more tokens to maximise performance, especially at higher effort levels [4]. Same rate, more units per unit of work.
The config line that proves Google knows this sits in the same post: use lower effort levels to minimise token overhead, or stay on 3.7 Flash, which remains fully supported for efficiency-first workloads [5]. A vendor pointing you at last month's model for cost reasons is doing you a favour. It is also telling you what the new one costs.
Then the date. Both halves of the introductory price double: input from $0.75 to $1.50, output from $3.75 to $7.50 per million tokens, exactly 2.0x each [2][1]. A unit of work that reads a million input tokens and writes a million output tokens costs $4.50 now and $9.00 after the expiry [2]. If the extra diligence lifts token use by half, the post-expiry bill is 3x what you are paying today [3]. That 50% is my illustration, not Google's number; nothing in the announcement quantifies the overhead. Google also says the release cadence is accelerated by long-running agentic loops that recursively evaluate and refine the models [19], which is a reason to expect another version before you finish migrating to this one.
On benchmarks, treat the table as a claim about someone else's repo. Google says 3.8 Flash matches Opus 5 on DeepSWE and beats GPT-5.6 Sol and Sonnet 5 [7]. For that to transfer, your tasks need to look like DeepSWE's: long-horizon engineering problems solved end to end with a verifiable finish. Where it still trails is computer use, 59% against 75.4% for Opus 5 on OSWorld-2.0 [12], and knowledge work, 1545 on GDPVal against 1824 for Opus 5, 279 points back [13][5]. The New Stack notes that Google's comparison set omits Anthropic's Fable 5.1, which shipped a day earlier [9]; on Terminal-Bench 4.0, Flash 3.8 scores 19.1% against Fable's 55.8% and Opus 5's 51.8% [8], a 2.9x gap to Fable [4]. On the coding-focused Terminal-Bench 2.1, Flash 3.8 beats its competitors [10], and The New Stack reads Terminal-Bench 4.0 as the outlier rather than the verdict [11]. The same publication points out that GLM-5.3, DeepSeek v4 Pro and Kimi K3 sit in the same league on DeepSWE 1.1 at better price/performance [20], which is the comparison the price change actually invites.
Fairwind is the more interesting piece of engineering discipline. Google says it invested in vulnerability fixing from the start and prioritised it over offensive capabilities like exploitation [14], and reports 47.2% pass@1 on Collinear's CWE-Bench against 47.8% for a leading frontier model at significantly lower cost [15]. Near-parity on patching at Flash cost, still under 50% first-try, means a human reads every diff. Access runs through about 650 trusted partners including CrowdStrike, Palo Alto Networks, Datadog, Snowflake, Wiz, Accenture and the Center for Internet Security [6]. The 3.5 Cyber equivalent was unnamed and, by The New Stack's reading, ad hoc [16]. Naming the programme and putting 650 partners behind it turns a capability into a procurement category.
Ranked by verification strength, evidence, and original report placement.
Google launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on Wednesday; it is the company's third Flash release in six weeks, and Gemini 3.7 Flash is only three weeks old.
Gemini 3.8 Flash is available at $0.75 per million input tokens and $3.75 per million output tokens, the same introductory price as 3.7 Flash; the introductory price expires on December 31, 2026 and then goes to $1.50/$7.50.
Google says 3.8 Flash "works harder": on complex tasks it exhibits greater diligence, executing extra reasoning steps and calling tools iteratively, and at times may use more tokens to maximise performance, especially at higher effort levels.
Gemini 3.8 Flash Cyber is available only to trusted defenders through Google's new Fairwind Program, comprising about 650 trusted partners including Accenture, CrowdStrike, the Center for Internet Security, Datadog, Palo Alto Networks, Snowflake and Wiz.
Google says 3.8 Flash often outperforms GPT-5.6 Sol, Claude Sonnet 5 and Opus 5 on complex engineering tasks, and on DeepSWE it matches Opus 5 while beating GPT-5.6 Sol and Sonnet 5.
On CWE-Bench, an external patching benchmark run by Collinear, Gemini 3.8 Flash Cyber records a pass@1 of 47.2% compared with 47.8% for a leading frontier model, at significantly lower cost.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 2, 2026
1 article · September 2, 2026
2 articles · September 2, 2026
2 articles · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Google's Opus comparison for Gemini 3.8 Flash ran entirely inside its own coding tool1 distinct publisher
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
security
Frontier labs put their best vulnerability-hunting models behind vetted-defender lists1 distinct publisher
build
Four frontier models in four days, and the cheapest number in your agent plan has an expiry date1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Vendor scoreboard, two outside checks
Nearly every capability number here originates in Google's own launch post, including the internal 20-language benchmark and the Chrome and Wiz results, which no one has replicated. Two figures escape that: Collinear's CWE-Bench run and the Artificial Analysis cost-per-task series The Decoder cites. The story's own load — the pricing schedule — rests on trade reporting rather than the announcement, since Google's text marks the rate 'introductory' with a footnote and never prints the date; The New Stack and The Decoder agree on the step-up, which is why it holds.
Everywhere Google, nowhere else yet
Distribution on day one is genuinely broad — AI Studio, Antigravity, Android Studio, Gemini Enterprise, the Gemini app, Search AI Mode, Sheets — and the Cyber variant reaches roughly 650 vetted partners. But everything we can actually observe being used is inside Google: Chrome Security's patch counts, its Cloud Vulnerability Research find. Wiz's recall figures come to us through Google's post rather than Wiz. No customer has told us what running this in production costs or catches.
'Same low cost' does not survive the per-task view
The overstatement is narrow but real: Google sells 3.8 Flash at 'the same speed and low cost' as 3.7 while conceding in the next section that it burns more tokens, and Artificial Analysis puts that at roughly 40% more per task. 'Often approaching frontier performance' also sits awkwardly beside 19.1% on Terminal-Bench 4.0 and a 279-point GDPVal deficit to Opus 5. Our own headline earns a discount too — the doubling is a promotion ending, and even afterwards this model costs a third of Opus 5 per token.
Launch-day economics on both sides of the byline
Two of our three publishers are Google's own blog, and the material they publish is a sales document: unnamed 'leading frontier models', internally built benchmarks, and a footnoted price that omits its expiry. The Fairwind names carry their own interest — Wiz's flattering recall numbers arrive inside the vendor's post, and security firms in a gated programme have reason to be quoted approvingly. The trade outlets are working a same-day release cycle off one announcement, with The Decoder closing on a subscription pitch.
Firm on the rate card, soft on the outcomes
We would bet on the pricing schedule and the release facts: three publishers, two of them independent of Google, agree on the dates and figures. We would not bet on the capability picture. Terminal-Bench 2.1's claimed lead comes with no score attached, the CWE-Bench rival is anonymous, and the 70% internal result is unauditable. What the model actually costs a team over a quarter is the number that matters and the one nobody has yet.