Skip to content

Leadership2 publishers3 min readPublished

Gemini 4 Argon ties two rivals on a leading AI index while some Google staff doubt its coding

Google's Gemini 4 Argon tied GPT-6 Astra and Claude Fable 5.1 at 53 on the Artificial Analysis index, though some Google staff say its coding lags. Google disputes them, and until paid API access has a date, buyers cannot check the score on their own code.

The Board Room · Leadership desk

Illustration accompanying Gemini 4 Argon ties two rivals on a leading AI index while some Google staff doubt its coding

What happened

  • On the same index, Argon trails Anthropic's two newest models: Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56.
  • Two people familiar with Argon said it appeared affected by 'benchmaxxing', meaning engineering effort aimed at test scores more than at doing the user's job well.
  • Google said it would be inaccurate to say Argon underperforms in coding, and an employee cited 'large consensus' internally that it is at the frontier.
  • Only cybersecurity partners in Google's Fairwind program can use Argon for now; paid API customers and AI Ultra subscribers come next, with no date set.
  • Gemini 3.5 Pro, the frontier model Google promised for June, was postponed several times internally over poor performance and never released.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • contradiction The case against Argon rests on anonymous staff and the case for it on examples Google chose, so the published record cannot settle which account holds for outside code.
  • exposure Front-end design is the weakness one employee named, and Google's showcases are a video decoder and data-center memory, so app-building teams carry the most uncertainty.
  • cost Planning a migration around Argon's wider release ties a roadmap to a date Google has not set, after its June frontier date passed without a model.

Five points separate Argon from Claude Opus 5.5 on the index, and three from Claude Sonnet 5.5 [1]. Google's strongest benchmark claims are narrower. Argon shares the top score with GPT-6 Astra on CWE-bench, a test of finding and patching security vulnerabilities, and Google says it set a new high on a test of long-horizon engineering work [13].

The doubts inside Google concern what those numbers leave out. One employee with direct access singled out front-end design, the work that decides how an app or website looks and feels [5]. Edwin Chen, founder of the AI startup Surge AI, speaking about benchmarks generally, described how relying on them can lead a lab to produce code in a given programming language while paying less attention to whether the finished app is easy to use or well designed [10]. "An analogy would be, 'Oh yeah, my kid got a really good score on the SAT', but the SAT doesn't translate into real-world performance," Chen said. "It's an incredibly pernicious problem." [11] The incentive starts with buyers. Implicator, summarizing Bloomberg's reporting, says labs prioritize these scores because customers often judge models by benchmark rankings [12].

Google's evidence comes from its own systems. "We've seen strong performance up close as Googlers have put the model through its paces in recent weeks, with many relying on it for their hardest coding and research problems," said Tulsee Doshi, who leads Gemini products at DeepMind [8]. In the launch example, Argon agents replaced 32,000 lines of SIMD code in a Rust port of Google's libgav1 video decoder, and Google said the result ran 2.7 times faster than the earlier Rust version with identical output [9]. Google also said the model helped its engineers free more than 300 tebibytes of memory across its data centers without new hardware [14]. Koray Kavukcuoglu, who runs DeepMind day to day, told a conference audience last week: "In my mind, it's a certainty that we are always gonna be at the frontier." [18]

A skeptic of the Bloomberg account would note that the critics are anonymous [15]. Google's staff also disagree among themselves, some believing Anthropic and OpenAI are improving faster and others that Google has caught up [16]. The objection is fair, and it applies to Google's side as well. A decoder rewrite and a memory clean-up are performance jobs on code Google knows well and chose to publish [9][14]. Neither is front-end work, the area one employee named [5]. The reports cite no results on code from outside Google.

I'd separate this quarter's decision from next quarter's. This quarter most teams cannot run Argon at all [1], and Google's last promised frontier date passed with no model [17]. The choice open now is whether to build an evaluation set from a team's own tickets and repositories before access arrives. When API customers do get in, a team with that set can judge Argon on its own code. A team without one will be choosing on a 53 that some of Google's own staff say overstates its coding [4].

What to watch

  • A release date for paid API customers and Google AI Ultra subscribers, the point at which teams outside Fairwind can test Argon on their own code.
  • Results from outside customers on front-end and app-building tasks, the area where employees said Argon struggles.
  • Whether Google publishes coding results on codebases it does not own, beyond the libgav1 decoder and data-center memory examples.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories