Build1 distinct publisher3 min readUpdated
Grok 4.6, Gemini 3.7 Flash, DeepSeek V4 Pro and GLM-5.3 all chase agents that stay on task. The pricing underneath them is moving faster than the benchmarks.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Four labs shipped frontier models inside four days during the week of August 11 to 18, 2026, and every one of them was tuned for the same property: agents that stay on task across many steps [1]. The more consequential news sat below the model layer, where DeepSeek raised prices on V4 Pro immediately after taking it to general availability [4], and 2027 DRAM and HBM capacity is reportedly already sold out [7].
Start with what the capability gains actually look like. Grok 4.6, released by SpaceXAI on August 12 alongside the closing of its Cursor acquisition [2], is not a bigger base model. The lab held the foundation constant and spent the budget on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning inside agentic environments [8]. It ships with a 500,000-token context window and a new xhigh reasoning-effort level [9], at $2 per million input tokens, $0.50 cached, and $6 per million output below 200K prompt tokens, doubling to $4, $1 and $12 above that threshold [10]. It is generally available as grok-4.6, is the default in Grok Build, and lands in Cursor with doubled included usage for the first week [11].
Artificial Analysis scores it 61 on its Intelligence Index, five points above Grok 4.5 and tied with GPT-5.6 Sol Max for third [12], with an AA-Briefcase Elo of 1,577 against 1,574 for Claude Fable 5 Max [13]. The number operators should care about is the token accounting: Artificial Analysis reports Grok 4.6 finishing AA-Briefcase workloads in roughly 53 turns and about 0.5 billion input tokens, against roughly 103 turns and 2 billion for Claude Opus 5 Max [14] - about 1.9 times fewer turns and four times fewer input tokens [1]. Fewer turns means less re-read context per step, which compounds [15].
Then the caveats, which are real. The AA-Briefcase and GDPval-AA v2 wins sit inside published confidence intervals, making them statistical ties [16]. SpaceXAI's own comparison table omits Claude Opus 5, which tops the index at 63, two points clear [17][5]. On coding, Grok 4.6 posts 65.9% on DeepSWE v1.1 against 73% for GPT-5.6 Sol Max, and 26% on Terminal-Bench v3.0, nearly double its predecessor and still last among the listed frontier models [18]. Artificial Analysis puts it at $0.84 per completed task, less economical than GPT-5.6 Luna and GLM-5.2 [19].
Google's release makes the cost point sharper. Gemini 3.7 Flash arrived on August 13, 23 days after 3.6 Flash [3], with identical specs - a 1,048,576-token input window, 65,536-token output limit, March 2026 cutoff [23] - and DeepSWE v1.1 at 65.3% against 49.0% for 3.6 Flash [26], a 16.3-point gain in 23 days [4]. That puts a Flash-tier model within 0.6 points of Grok 4.6 on the same coding test [2] at $0.75 per million input tokens rather than $2.00, a 2.7x difference [20][3]. Except the $0.75 and $3.75 rates expire on December 31, 2026, after which they double to $1.50 and $7.50, exactly what 3.6 Flash cost at launch [21]. Google applied the promotional rate to 3.6 Flash as well, so until year-end the migration question is capability, not list price [22].
What to watch: January 1, when Gemini Flash pricing resets to double [21]; whether DeepSeek's post-GA increase [4] is the start of a pattern; and the memory market, since sold-out 2027 DRAM and HBM capacity [7] sets the floor under every per-token price in this paragraph. Also worth tracking is the MCP stateless spec, now in its adoption window [6], and whether Z.ai's GLM-5.3 cybersecurity claims, which CVE databases partially support [5], hold up under scrutiny.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
During the week of August 11 to 18, 2026, four labs shipped frontier models within four days of each other, and every one was tuned for agents that stay on task.
SpaceXAI released Grok 4.6 on August 12, 2026 and closed its Cursor acquisition.
Google released Gemini 3.7 Flash on August 13, 2026, 23 days after Gemini 3.6 Flash.
DeepSeek took V4 Pro to general availability and then raised its prices.
Grok 4.6 is a post-training upgrade over Grok 4.5 rather than a larger base model; the lab held the foundation constant and spent the improvement budget on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning inside agentic environments.
Grok 4.6 has a 500,000-token context window and a new xhigh reasoning-effort level above the existing ladder.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific but single-sourced
Claims are unusually concrete - model IDs, token prices, context limits, named benchmarks with figures - and the author separates independent Artificial Analysis results from vendor-reported Google numbers and flags confidence-interval ties and an omitted index leader. But everything comes from one secondhand weekly roundup with no primary release notes, pricing pages, or benchmark pages in the cluster, and the two most consequential assertions (2027 memory sell-out, GLM-5.3 security) carry no supporting detail at all.
Shipped and reachable, usage unknown
All four models are past announcement: Grok 4.6 is GA on the xAI API, default in Grok Build and live in Cursor; Gemini 3.7 Flash spans six Google surfaces; DeepSeek V4 Pro reached GA and was listed on OpenRouter the same day. That is real distribution. What is missing is any usage, traffic, customer, or migration disclosure - no seat counts, token volumes, or named production deployments - so adoption is credible at the availability layer only.
Slightly overstated at the vendor layer
The article itself deflates most hype - it labels bolded wins as statistical ties, notes SpaceXAI's table excludes the index leader, tags Google's headline numbers as vendor-reported, and shows Grok last on Terminal-Bench and worse than cheaper rivals on cost per task. The residual gap is upward: 'four frontier models in four days, all tuned for agents' is framed as a step change while a Flash-class model lands within 0.6 points of the frontier release on DeepSWE, and the framing-heavy claims about GLM-5.3 security and 2027 memory sell-out carry no evidence.
Vendor pricing and benchmark incentives visible
The underlying material is largely vendor-controlled and vendor-timed. Google's half-price Flash rate is promotional with a hard December 31, 2026 expiry that returns pricing to the prior launch level, and the promo was extended to the older model to steer migration on capability rather than price. SpaceXAI's own comparison table excludes the model that currently leads the index, and its Cursor rollout doubles included usage for one week inside a product it just acquired. DeepSeek raised prices through a silent pricing-page edit with no announcement. Each is a documented incentive to shape perception or lock in usage.
Moderate-low: one publisher, dense specifics
Confidence is limited chiefly by cluster structure: a single dev.to roundup, no primary or corroborating publisher, and a truncated body that omits the DeepSeek pricing detail the summary asserts. Confidence is lifted by the density and checkability of the numbers, explicit attribution to Artificial Analysis versus vendor reporting, and the author's own caveats - but the pricing, benchmark, and availability facts remain uncorroborated within this cluster.
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
invest
Speed becomes a SKU: OpenAI and Google put a separate price on latency3 distinct publishers
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
GLM-5.3 kept the base model and bought ten times the environments instead2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 18, 2026