Skip to content

BuildIndependently confirmed3 publishers2 min readPublished

Google tests a coding-tuned Gemini 4 checkpoint its staff liken to Anthropic's Opus 5.5

Google is internally testing a Gemini 4 checkpoint called Carbon that one employee compared to Anthropic's Opus 5.5 for coding, Business Insider reported. The model is unreleased, and the comparison rests on early internal feedback.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Google tests a coding-tuned Gemini 4 checkpoint its staff liken to Anthropic's Opus 5.5
Generated illustration

What happened

  • Carbon was deployed over recent days on Jetski, Google's internal coding platform, and reportedly beats the Argon checkpoint mainly on programming tasks.
  • Google announced Argon on September 30, opening access first to selected cybersecurity partners, with paid API customers and Google AI Ultra subscribers to follow.
  • In Antigravity's model selector, Argon offers three context sizes: 256K by default, 512K at about 1.3x quota per turn, and 900K at about 1.8x quota per turn.
  • Business Insider identified the public Argon as a checkpoint called Barium-B, with Carbon a separate iteration described as potentially more capable.

Why it matters

  • decision A team that standardized on one coding model has a reason to keep its own evaluation harness current, because a credible second coding model is now in testing.
  • cost Argon's context sizes are priced in quota multipliers, so teams that lean on long context for agent runs pay more per turn and have to budget context the way they budget tokens.
  • constraint Google is shipping Argon, not Carbon, and an employee likened early Argon builds to the older Opus 5, so the model customers get may trail the internal checkpoint that drew the newer comparison.

Carbon is a checkpoint inside the Gemini 4 lineup, not a separate product tier [16]. Google maps its models to jobs, writing that it uses "Argon for frontier reasoning, Flash for speed and volume, Omni for generative media, and Gemma for lightweight, open-weights edge workloads" [9]. Argon is the top reasoning tier, set against Anthropic's Opus and OpenAI's Astra [11]. In internal documents the shipping Argon was called Barium-B, and one employee referred to Carbon as the "Gemini pro next model," which suggests it would ship under the Argon name [3][17].

The Opus 5.5 comparison came out of hands-on testing, not a published benchmark. Google is running Carbon through Jetski, its internal name for the platform behind the Antigravity agent tool [10]. For a figure like that to transfer to another team, it has to hold up on a benchmark that team can run against its own code.

Google is iterating fast. Google DeepMind's Vedant Misra answered the Business Insider report on X with "Have you heard of recursive self improvement," a reference to using AI to build better AI [13]. OpenAI and Anthropic have described similar work [19]. Logan Kilpatrick has said the team is still working on getting the most out of Argon [14].

Both outlets trace the comparison to the same Business Insider report [5][6]. Screenshots and internal chats are thinner evidence than a model you can call and benchmark yourself [5].

What to watch

  • Whether Carbon ships as an update to Argon or as its own Gemini 4 model, and under what name.
  • A public, reproducible coding benchmark for Argon or Carbon against Anthropic's Opus 5.5.
  • Argon's wider rollout to paid API customers and AI Ultra subscribers, which Google has not dated but could begin as soon as next week.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence35
Adoption8
Hype gap+45
Incentives
Insufficient
Confidence45

Perspective Coverage

3 publishers
Builder
Builder 43%
Operator
Operator 25%
Investor
Investor 32%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    One Google employee compared the Carbon checkpoint to Anthropic's strongest coding model, Opus 5.5, though it still needs more testing.

    ReportedSupportedSource: Business Insider, via one Google employee, as reported by The Decoder3 sources— create a free account to open themView cited source
  2. [2]

    Carbon was deployed on Google's internal coding platform Jetski over the past few days and is said to beat Argon mainly on programming tasks.

  3. [3]

    Internal documents show the now-unveiled Argon was previously called "Barium-B."

Sources

3 independent publishers whose own reporting we read for this story.

  1. businessinsider.com

    1 article · October 9, 2026

    Google is about to roll out a new AI model. Employees say they're testing another that's way better.
  2. testingcatalog.com

    1 article · October 9, 2026

    Gemini 4 Argon hints emerge as Google tests Carbon checkpoint
  3. the-decoder.com

    1 article · October 10, 2026

    Google's Gemini 4 "Carbon" model is reportedly matching Anthropic's Opus 5.5 coding performance

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories