Skip to content

Build1 publisher3 min readPublished

Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.

The Engineer · Build desk

Illustration accompanying Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

What happened

  • On DeepSWE, Flash spends 166 steps and 143,000 output tokens per task, 2.7 times the steps and 2.4 times the output tokens of GPT-5.6 Sol.
  • Artificial Analysis ranks Flash 28th of 202 models on its Intelligence Index, a run that took 140 million output tokens against a median of 79 million.
  • Meta shipped Muse Spark 1.3 four hours after Flash, with a contributor endpoint at $0.10 in and $0.20 out per million tokens where Meta uses your data to improve its products.
  • Google keeps Gemini 3.7 Flash fully supported for what it calls efficiency-first workloads.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Agent budgets set at Flash's September rate cover about half of the January bill for the same work, unless steps per task fall.
  • decision Choosing a Muse Spark 1.3 tier becomes a per-workload data-governance call, since permission to train on sessions is the only thing separating the two prices.
  • constraint Flash's per-task cost transfers only to agent loops that resemble mini-swe-agent; a harness that caps steps would change the bill and plausibly the score.
  • constraint Chat interfaces feel Flash's 12.74-second time to first token, against a 3.33-second median, far more than background agents do.

Google's launch post explains the gains in four words: "3.8 Flash works harder." [7] According to the post, it works harder by "executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance." [7] A per-token rate cannot tell a buyer how many tokens a task will take.

The DeepSWE v1.1 board, as reported in a dev.to comparison of the launches, measures that directly. Datacurve runs 113 tasks through the same mini-swe-agent harness for every model [6], so step and token counts compare across models on that board. For the same 74%, Opus 5 costs about five times as much per task as Flash [4][2]. Google deserves credit for this. Flash spends many cheap tokens and still finishes the task for less. Artificial Analysis saw the same spending pattern on a different workload. Its Intelligence Index run on Flash used about 1.8 times the median output tokens [6], and grading it cost $1,077.95 [8].

Google's footnote sets the price change: "Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply." [3] The introductory rate lasts about four months from the September 2 launch [1][7]. At Flash's DeepSWE average, output tokens alone cost about $0.54 per task at today's $3.75 per million and about $1.07 at January's $7.50 [2][1]. Input tokens are excluded from those figures, so the full cost is higher. If token use per task holds, and the $2.36 was priced at the introductory rate, a full DeepSWE task goes to about $4.72 [3]. Opus 5's $11.84 is still about 2.5 times that [5].

Before trusting the $2.36 or the $4.72, I'd log steps and output tokens per finished task on my own harness, against my own tasks. A leaderboard's per-task figure is a measurement of Datacurve's workload.

Meta's Muse Spark 1.3 complicates the per-token view further. Its standard output rate of $4.25 per million sits above Flash's $3.75 today and below Flash's $7.50 from January [10][8]. The per-token ranking of the two models flips on a calendar date. The contributor endpoint is about 21 times cheaper than standard on output and 12.5 times cheaper on input [4]. The comparison does not include a DeepSWE result for Muse Spark 1.3, so its steps per task are unmeasured.

On Hacker News, user hiddencost wrote of Google's pricing: "you're effectively planning to charge users twice as much for a model that is no longer frontier." [12] The thread's top comment was Simon Willison's quick test: "make me a cool thing in html" returned a particle simulation in 13 seconds for 1.8 cents [13]. Another commenter noticed that its "60 FPS" counter was hard-coded into the page [14].

What to watch

  • Whether Google extends or revises the Gemini 3.8 Flash introductory price before it expires on December 31, 2026.
  • A DeepSWE or Artificial Analysis per-task result for Muse Spark 1.3, which would allow a cost-per-finished-task comparison with Flash.
  • Whether the next Flash release cuts steps per task, given Google shipped three Flash models in six weeks.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories