Skip to content

Build1 publisher3 min readPublished

Google is testing Gemini 4 inside Antigravity, its agentic coding platform, before release

Google DeepMind's Koray Kavukcuoglu says Gemini 4 is in post-training and under test in Antigravity, with release hoped for well before the end of 2026. With no benchmarks published, the first evidence teams get will come from their own agent tasks.

The Engineer · Build desk

Illustration accompanying Google is testing Gemini 4 inside Antigravity, its agentic coding platform, before release

What happened

  • Google DeepMind chief Koray Kavukcuoglu said on September 23 that Gemini 4 is in post-training and being tested internally in Antigravity, Google's agentic software-development platform.
  • He said Google wants an early post-training version out "as soon as possible" and hopes to make Gemini 4 available well before the end of 2026, but gave no launch date.
  • The post-training work under way includes refining the model's behavior, running safety tests and adding guardrails before wider availability.
  • Kavukcuoglu said Google had "taken a little bit of a step back" from Gemini 3.5 Pro to focus on Flash models, and Google has not said whether 3.5 Pro will ship.
  • Sundar Pichai said on Google's July 22 earnings call that Antigravity had more than 2.4 million weekly active users.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision The first usable evidence on Gemini 4 for agent work will come from teams' own task suites, so an evaluation harness built on real repositories needs to exist before the model arrives.
  • constraint "Well before the end of 2026" is Kavukcuoglu's hope, so migration budgets and cutover plans cannot yet be pinned to a quarter.
  • decision Roadmaps that assumed a Gemini 3.5 Pro need a fallback to current Flash models or to Gemini 4, since Google has moved its flagship effort to Gemini 4.

"The conversation is more about are we able to build intelligent agents that we can trust," Kavukcuoglu said, in a reproduction of The Information interview that The Decoder corroborated [3]. Runtimewire argues the Antigravity test puts that question ahead of the usual frontier-model scorekeeping: can the next flagship power agents people will trust to carry out work [5]. I agree with that ordering for teams that run agents. A coding agent can go wrong at any step of a long task, and a one-shot benchmark score grades one answer.

An internal run in Antigravity tells Google how Gemini 4 behaves in a software workflow Google controls [7]. According to Runtimewire, it does not establish that the model will beat competitors at coding, or that outside developers will get it through Antigravity first [7]. Google has not published benchmark results or a release format [8]. For Google's internal findings to predict anyone else's results, that team's repositories, tool permissions and task lengths would have to resemble the work Google's engineers give the agent.

The early release plan changes how testing has to be done. Runtimewire describes an early post-training version as a way for Google to keep iterating on a model already past initial training while tuning and testing continue [9]. The same report calls the plan a stated intention, not a schedule or a promise of general availability [9]. The checkpoint teams test first can differ from the one Google later calls finished. A pass rate on the early version measures the early version. The harness needs to log the exact model identifier with every run, so the suite can be repeated when Google swaps it.

Google's naming and release sequence has made the flagship timeline harder to read, in Runtimewire's account [17]. Gemini 3 Pro arrived in November 2025 and Gemini 3.1 Pro followed in February 2026 [11], about three months apart [1]. Since then Google has issued several Flash updates, most recently Gemini 3.8 Flash [12]. The Flash numbering has now passed a Pro version that may never ship [6].

Google already has developer products where a stronger model could be used [14]. Pichai said on the July 22 earnings call that more than 9 million developers were building monthly with Google's models and that its model APIs were processing about 22 billion tokens per minute [13]. Those are company-reported figures for the existing business, not Gemini 4 adoption [14].

Kavukcuoglu said Google remains confident it will stay at the frontier, despite questions about whether it has fallen behind OpenAI and Anthropic [15]. That claim will be judged once Gemini 4 can be compared and used outside Google's internal systems, according to Runtimewire [16].

What to watch

  • Google publishing Gemini 4 benchmark results and a release format, including whether outside developers get it through the API, Antigravity, or both.
  • A dated release of the early post-training version, and whether its model identifier changes as tuning and guardrail work continue.
  • Whether Google ships or drops Gemini 3.5 Pro.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories