Skip to content

Build1 publisher3 min readPublished

Four frontier models in four days, and the cheapest number in your agent plan has an expiry date

Grok 4.6, Gemini 3.7 Flash, DeepSeek V4 Pro and GLM-5.3 all chase agents that stay on task. The pricing underneath them is moving faster than the benchmarks.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Four frontier models in four days, and the cheapest number in your agent plan has an expiry date
Generated illustration

What happened

  • During the week of August 11 to 18, 2026, four labs shipped frontier models within four days of each other, and every one was tuned for agents that stay on task.
  • SpaceXAI released Grok 4.6 on August 12, 2026 and closed its Cursor acquisition.
  • Google released Gemini 3.7 Flash on August 13, 2026, 23 days after Gemini 3.6 Flash.
  • DeepSeek took V4 Pro to general availability and then raised its prices.
  • Z.ai announced GLM-5.3 with cybersecurity claims that real CVE databases partially back up.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Four labs shipped frontier models inside four days during the week of August 11 to 18, 2026, and every one of them was tuned for the same property: agents that stay on task across many steps [1]. The more consequential news sat below the model layer, where DeepSeek raised prices on V4 Pro immediately after taking it to general availability [4], and 2027 DRAM and HBM capacity is reportedly already sold out [7].

Start with what the capability gains actually look like. Grok 4.6, released by SpaceXAI on August 12 alongside the closing of its Cursor acquisition [2], is not a bigger base model. The lab held the foundation constant and spent the budget on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning inside agentic environments [8]. It ships with a 500,000-token context window and a new xhigh reasoning-effort level [9], at $2 per million input tokens, $0.50 cached, and $6 per million output below 200K prompt tokens, doubling to $4, $1 and $12 above that threshold [10]. It is generally available as grok-4.6, is the default in Grok Build, and lands in Cursor with doubled included usage for the first week [11].

Artificial Analysis scores it 61 on its Intelligence Index, five points above Grok 4.5 and tied with GPT-5.6 Sol Max for third [12], with an AA-Briefcase Elo of 1,577 against 1,574 for Claude Fable 5 Max [13]. The number operators should care about is the token accounting: Artificial Analysis reports Grok 4.6 finishing AA-Briefcase workloads in roughly 53 turns and about 0.5 billion input tokens, against roughly 103 turns and 2 billion for Claude Opus 5 Max [14] - about 1.9 times fewer turns and four times fewer input tokens [1]. Fewer turns means less re-read context per step, which compounds [15].

Then the caveats, which are real. The AA-Briefcase and GDPval-AA v2 wins sit inside published confidence intervals, making them statistical ties [16]. SpaceXAI's own comparison table omits Claude Opus 5, which tops the index at 63, two points clear [17][5]. On coding, Grok 4.6 posts 65.9% on DeepSWE v1.1 against 73% for GPT-5.6 Sol Max, and 26% on Terminal-Bench v3.0, nearly double its predecessor and still last among the listed frontier models [18]. Artificial Analysis puts it at $0.84 per completed task, less economical than GPT-5.6 Luna and GLM-5.2 [19].

Google's release makes the cost point sharper. Gemini 3.7 Flash arrived on August 13, 23 days after 3.6 Flash [3], with identical specs - a 1,048,576-token input window, 65,536-token output limit, March 2026 cutoff [23] - and DeepSWE v1.1 at 65.3% against 49.0% for 3.6 Flash [26], a 16.3-point gain in 23 days [4]. That puts a Flash-tier model within 0.6 points of Grok 4.6 on the same coding test [2] at $0.75 per million input tokens rather than $2.00, a 2.7x difference [20][3]. Except the $0.75 and $3.75 rates expire on December 31, 2026, after which they double to $1.50 and $7.50, exactly what 3.6 Flash cost at launch [21]. Google applied the promotional rate to 3.6 Flash as well, so until year-end the migration question is capability, not list price [22].

What to watch: January 1, when Gemini Flash pricing resets to double [21]; whether DeepSeek's post-GA increase [4] is the start of a pattern; and the memory market, since sold-out 2027 DRAM and HBM capacity [7] sets the floor under every per-token price in this paragraph. Also worth tracking is the MCP stateless spec, now in its adoption window [6], and whether Z.ai's GLM-5.3 cybersecurity claims, which CVE databases partially support [5], hold up under scrutiny.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories