BuildIndependently confirmed3 publishers2 min readPublished
Google tests a coding-tuned Gemini 4 checkpoint its staff liken to Anthropic's Opus 5.5
Google is internally testing a Gemini 4 checkpoint called Carbon that one employee compared to Anthropic's Opus 5.5 for coding, Business Insider reported. The model is unreleased, and the comparison rests on early internal feedback.
The Engineer · Build desk

What happened
- Carbon was deployed over recent days on Jetski, Google's internal coding platform, and reportedly beats the Argon checkpoint mainly on programming tasks.
- Google announced Argon on September 30, opening access first to selected cybersecurity partners, with paid API customers and Google AI Ultra subscribers to follow.
- In Antigravity's model selector, Argon offers three context sizes: 256K by default, 512K at about 1.3x quota per turn, and 900K at about 1.8x quota per turn.
- Business Insider identified the public Argon as a checkpoint called Barium-B, with Carbon a separate iteration described as potentially more capable.
Why it matters
- decision A team that standardized on one coding model has a reason to keep its own evaluation harness current, because a credible second coding model is now in testing.
- cost Argon's context sizes are priced in quota multipliers, so teams that lean on long context for agent runs pay more per turn and have to budget context the way they budget tokens.
- constraint Google is shipping Argon, not Carbon, and an employee likened early Argon builds to the older Opus 5, so the model customers get may trail the internal checkpoint that drew the newer comparison.
Carbon is a checkpoint inside the Gemini 4 lineup, not a separate product tier [16]. Google maps its models to jobs, writing that it uses "Argon for frontier reasoning, Flash for speed and volume, Omni for generative media, and Gemma for lightweight, open-weights edge workloads" [9]. Argon is the top reasoning tier, set against Anthropic's Opus and OpenAI's Astra [11]. In internal documents the shipping Argon was called Barium-B, and one employee referred to Carbon as the "Gemini pro next model," which suggests it would ship under the Argon name [3][17].
The Opus 5.5 comparison came out of hands-on testing, not a published benchmark. Google is running Carbon through Jetski, its internal name for the platform behind the Antigravity agent tool [10]. For a figure like that to transfer to another team, it has to hold up on a benchmark that team can run against its own code.
Google is iterating fast. Google DeepMind's Vedant Misra answered the Business Insider report on X with "Have you heard of recursive self improvement," a reference to using AI to build better AI [13]. OpenAI and Anthropic have described similar work [19]. Logan Kilpatrick has said the team is still working on getting the most out of Argon [14].
Both outlets trace the comparison to the same Business Insider report [5][6]. Screenshots and internal chats are thinner evidence than a model you can call and benchmark yourself [5].
What to watch
- Whether Carbon ships as an update to Argon or as its own Gemini 4 model, and under what name.
- A public, reproducible coding benchmark for Argon or Carbon against Anthropic's Opus 5.5.
- Argon's wider rollout to paid API customers and AI Ultra subscribers, which Google has not dated but could begin as soon as next week.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption8
- Hype gap+45
- Incentives
- Insufficient
- Confidence45
Perspective Coverage
3 publishers- Builder
- Builder 43%
- Operator
- Operator 25%
- Investor
- Investor 32%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
One Google employee compared the Carbon checkpoint to Anthropic's strongest coding model, Opus 5.5, though it still needs more testing.
ReportedSupportedSource: Business Insider, via one Google employee, as reported by The Decoder3 sources— create a free account to open themView cited source - [2]
Carbon was deployed on Google's internal coding platform Jetski over the past few days and is said to beat Argon mainly on programming tasks.
ReportedSupportedSource: Business Insider3 sources— create a free account to open themView cited source - [3]
Internal documents show the now-unveiled Argon was previously called "Barium-B."
ReportedSupportedSource: Business Insider3 sources— create a free account to open themView cited source - [4]
Whether Carbon ships as an Argon update or a standalone model remains unclear.
- [5]
According to documents, screenshots, and internal chats seen by Business Insider, Google is testing several Gemini 4 variants called Argon, Barium, and Carbon.
ReportedSupportedSource: Business Insider2 sources— create a free account to open themView cited source - [6]
According to early employee feedback reported by Business Insider, the Carbon checkpoint reportedly feels comparable to Claude Opus 5.5 for coding.
ReportedSupportedSource: Business Insider, via TestingCatalog3 sources— create a free account to open themView cited source - [7]
Early Argon versions reminded another employee of the older Opus 5 on some coding tasks, even though internal feedback was positive overall.
ReportedSupportedSource: Business Insider2 sources— create a free account to open themView cited source - [8]
The report identifies Barium-B as the checkpoint selected for the public Argon release, while Carbon represents a separate and potentially more capable iteration.
ReportedSupportedSource: Business Insider, via TestingCatalog3 sources— create a free account to open themView cited source - [9]
Google writes it uses "Argon for frontier reasoning, Flash for speed and volume, Omni for generative media, and Gemma for lightweight, open-weights edge workloads."
- [10]
Carbon is reportedly being tested through Jetski, Google's internal name associated with Antigravity.
- [11]
Argon is Google's most capable reasoning model, comparable to Anthropic's Opus or OpenAI's Astra.
- [12]
Google officially announced Argon on September 30, initially granting access to selected cybersecurity partners, with paid API customers and Google AI Ultra subscribers expected to follow.
- [13]
Google DeepMind employee Vedant Misra responded to the Business Insider report on X by writing, "Have you heard of recursive self improvement."
- [14]
Google's Logan Kilpatrick confirmed the team is working on getting the most out of Argon.
- [15]
Antigravity's model selector references Argon with three context configurations: 256K by default, 512K consuming approximately 1.3x quota per turn, and 900K consuming around 1.8x quota per turn.
- [16]
Carbon and Barium are likely checkpoints or updates within the Argon family, not separate tiers.
- [17]
One employee internally called Carbon the "Gemini pro next model," suggesting it would ship under the Argon name.
ReportedContestedSource: Business Insider3 sources— create a free account to open themView cited source - [18]
Argon's wider rollout could begin as soon as next week, although Google has not confirmed a date.
- [19]
OpenAI and Anthropic have reported similar progress on using AI to build new models.
Sources
3 independent publishers whose own reporting we read for this story.
- businessinsider.comGoogle is about to roll out a new AI model. Employees say they're testing another that's way better.
1 article · October 9, 2026
- testingcatalog.comGemini 4 Argon hints emerge as Google tests Carbon checkpoint
1 article · October 9, 2026
- the-decoder.comGoogle's Gemini 4 "Carbon" model is reportedly matching Anthropic's Opus 5.5 coding performance
1 article · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Entities
- OpenAIFollow
- Claude Opus 5.5Follow
- GoogleFollow
- Logan KilpatrickFollow
- Vedant MisraFollow
- Google DeepMindFollow
- CarbonFollow
- AntigravityFollow
- Gemini 4Follow
- JetskiFollow
- Gemini 4 ArgonFollow
- AnthropicFollow