Leadership1 distinct publisher3 min readPublished
Google held its token rates flat, but the measured cost of an Intelligence Index task rose from $0.40 to $0.58 while the score still trails Fable and Sol. Picking a model for agent work is now a pricing exercise.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The cleanest way to read this release is as a price per marginal index point. Gemini 3.7 Flash delivered 56 at high effort for a measured $0.40 a task, and 3.8 delivers 59 for $0.58 [11][10][2]. The increment works out to $0.18 for three points, or $0.06 a point, against roughly $0.0071 a point that 3.7 averaged across its whole score [3][4]. Marginal intelligence at this end of the curve costs about eight times what the average point costs [5], and a team that books a version bump as a refresh will find a purchase on the invoice.
The rate card explains none of it. Google held the introductory rates at $0.75 per million input tokens and $3.75 per million output through December 31, 2026, and did not change them between the two versions [8][9]. What changed was consumption: output across the evaluation went from 64 million tokens to 120 million, up 87.5 percent [12][2]. Google's model card names the mechanism without dressing it up: "At times, the model might use more tokens to maximize performance, especially at higher effort levels" [19]. Artificial Analysis's per-task figures are weighted averages rather than provider list prices, and two models on identical per-token rates can produce different bills on the same evaluation [18].
The same arithmetic applied upward is what makes this a purchasing question rather than a ranking. Sol's two extra points over Gemini cost about $0.19 each, and Fable's seven cost about $0.44 each [6][7]. Fable's number is also not a clean single-model reading, because its default fallback means the run tested a system [17]. Below Gemini, GLM-5.3 Flash gives up two points at about a sixth of the per-task cost [6][7]. That is four prices for a unit of measured intelligence, and the index offers no basis for deciding which unit a given workload can convert into finished work.
Speed is the one axis where Gemini leads outright in this snapshot, at 304.6 output tokens a second against Sol's 72.4, Fable's 66.4 and GLM's 42.8 [13][14][15][16]. On the full run that is roughly 109 hours of generation for Gemini against about 269 for Sol [9][10], even though Gemini emitted 71 percent more output tokens [11]. For agent work that loops through many turns before returning anything, throughput is the cost a user feels, and it does not appear in the per-task price.
On a $0.58 task, 45 percent looks like rounding error for this quarter [1], and that reading holds up for now. The date that changes it is January 1, 2027, when the introductory rates double [8]. Hold token use where it is now and the same task costs $1.16, which is above what Sol measures today [8]. A team standardizing on 3.8 for agent work is accepting a per-task cost that reprices once, on a known date, by a known multiple.
Cheaper than the two systems above it, faster than everything in the snapshot: that is the board-deck version, and it is missing two things. The comparison ran at high effort while Google's default is medium, so the numbers most teams would actually generate are absent from the record [21]. And the aggregate disagrees with the domains. On Finance Agent v2, 3.8 led three of nine categories while 3.7 still led general quantitative analysis [22], and on the held-out Harvey legal test 3.8 resolved 10.0 percent of complete tasks against Muse Spark 1.2's 25.42 percent [23]. We do not know from this snapshot what 3.8 costs at the setting most production traffic will run on, and that is the number a buyer needs.
Ranked by verification strength, evidence, and original report placement.
Google introduced Gemini 3.8 Flash and the access-restricted Gemini 3.8 Flash Cyber on September 2.
On the Artificial Analysis Intelligence Index snapshot accessed September 2, Gemini 3.8 Flash at high effort scored 59.
GPT-5.6 Sol at maximum effort scored 61 and cost $0.95 per Intelligence Index task in the same snapshot.
Claude Fable 5.1 at maximum adaptive effort scored 66 and cost $3.69 per Intelligence Index task in the same snapshot.
Gemini 3.8 Flash's measured cost was $0.58 per Intelligence Index task.
Open-weights GLM-5.3 Flash scored 57, two points below Gemini 3.8 Flash, at $0.09 per task and $138.02 for the full run.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Gemini 3.8 Flash's introductory price doubles on December 31, 20265 distinct publishers
leadership
Meta's Muse Spark 1.3 matches three flagship models at 55 cents a task1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
product
Google's new Flash buys its benchmark wins with extra tokens1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise figures, one reader, one afternoon
Everything load-carrying traces to outside scoreboards — Artificial Analysis, CWE-Bench, Finance Agent, Harvey — read on September 2 by implicator.ai, which does the arithmetic in the open and marks its own soft spots: weighted averages instead of list prices, Fable's fallback muddying its 66, high effort standing in for a default the model does not use. What is absent is a second reading. Google's cyber numbers, the only ones flattering the gated model, come with the test sets and competitor names withheld, so they cannot be checked at all.
Shipped and scored, but nobody's usage is visible
We can confirm existence and availability: two models dated September 2, one callable at published rates, one behind Fairwind approval with named-user controls and no resale. Beyond that, silence. No developer counts, no workloads moved off 3.7, no sense of how many organisations cleared Fairwind. Appearing on four leaderboards is evidence that a benchmark house ran the model, not that anyone is running it in production.
The overstatement sits upstream of the reporting
implicator.ai runs cold — its own headline is the cost increase, not the capability. The stretch belongs to the launch it covers. Holding the token rate flat invites the reading that cost held flat, and the measured bill rose 45% anyway because the model talks more at high effort. The cyber variant is claimed to cost "significantly" less while the leaderboard that prices Fable 5 at $10.27 and Sol at $2.29 lists nothing for it. Modest positive gap, and it is the vendor's, not the desk's.
Vendor supplies the flattering half; the desk sells subscriptions
Two pulls, both worth naming. Google authored the model card, the launch framing and the only cyber figures that favour its model, while declining to name the test sets or the larger commercial systems it says it beat 2.6-to-1 on Chrome patches. And this reporting arrives wrapped in a morning-briefing signup, which rewards a sharp, quotable cost delta over a hedged one. Neither pressure touches the arithmetic, since the scoreboards are outside parties — but what got measured on September 2, and at which effort level, was somebody's choice.
Solid on the numbers, thin on their shelf life
The specific figures deserve trust: they are checkable, internally consistent, and the arithmetic on top of them reproduces. What lowers the reading is scope. One publisher, one access date, one effort setting, and a leaderboard that gets revised. Ask whether 3.8 Flash is still 45% dearer per task at default medium effort in a month and this reporting cannot answer.