Google is internally testing a new Gemini 4 checkpoint codenamed Carbon on its Jetski coding platform, with at least one employee describing its coding performance as comparable to Anthropic's Claude Opus 5.5 - though Google has confirmed no benchmarks, release date, or public availability for Carbon.
TECH

Google is internally testing a new Gemini 4 checkpoint codenamed Carbon on its Jetski coding platform, with at least one employee describing its coding performance as comparable to Anthropic's Claude Opus 5.5 - though Google has confirmed no benchmarks, release date, or public availability for Carbon.

31+
Signals

Strategic Overview

  • 01.
    Google employees have spent the past few days internally testing a new Gemini 4 checkpoint codenamed Carbon on Jetski, Google's internal coding platform, with at least one tester saying its coding performance feels comparable to Anthropic's Claude Opus 5.5.
  • 02.
    Carbon follows Argon, the publicly announced Gemini 4 flagship that began rolling out around September 30-October 1, 2026; internally, Argon corresponds to a checkpoint called Barium-B, and Carbon represents a later checkpoint still confined to internal testing.
  • 03.
    Google has not confirmed Carbon's benchmarks, release date, or whether it will ship as an Argon update, a separate Gemini 4 model, or stay internal-only; the checkpoint has no public model ID, API endpoint, pricing, or model card, and doesn't appear in AI Studio, Antigravity, or Vertex AI.
  • 04.
    Some employees attribute Carbon's coding gains to Recursive Self-Improvement (RSI) techniques, though this remains an internal rumor without official Google confirmation.

Deep Analysis

A Leap Confirmed by a Single Anonymous Screenshot

The entire Carbon story rests on internal Slack-style messages and screenshots relayed to a single Business Insider reporter, not on anything Google has confirmed [1]. The checkpoint has no public model ID, no API endpoint, no published price, no model card, and no stated evaluation methodology - it doesn't even appear in the menus for AI Studio, Antigravity, or Vertex AI [2]. The much-repeated 'feels like Opus 5.5' comparison is one employee's subjective impression, explicitly hedged with a caveat that more testing is needed [3], and online reaction has already needled the reporting itself, questioning whether the outlet understood what a model checkpoint even is [1].

From Argon to Barium-B to Carbon: Decoding Google's Checkpoint Pipeline

Google's naming makes more sense once you see the lineage: the publicly announced Gemini 4 Argon, which began rolling out around September 30, 2026, corresponds internally to a checkpoint called Barium-B, while Carbon is described as a later checkpoint still confined to internal testing [1][2]. That pipeline traces back to Google Antigravity, the agentic coding environment launched alongside Gemini 3 in November 2025 and now tied to the internal Jetski platform where Carbon is being run [4]. By September 23, 2026, DeepMind's Koray Kavukcuoglu had already said Gemini 4 was in post-training and being tested inside Antigravity, with hopes of an early post-training release before year-end [5].

The Opus 5.5 Benchmark Google Can't Officially Claim

Reports frame Carbon's internal testing as a direct response to Anthropic's Claude Opus 5.5, with Google racing to close or exceed the coding-performance gap just weeks after Argon's launch [3]. The urgency shows in how fast the iteration is moving: Carbon surfaced in internal tests before Argon had even finished its staged public rollout, which outlets have read as part of a broader push to win developer mindshare in agentic coding [6]. Some employees go further, attributing the jump to Recursive Self-Improvement (RSI) techniques - an unconfirmed internal theory, not an official Google explanation [7]. What Google can point to publicly is Argon's own benchmark record: 68% on CWE-bench v1 (tied for first place) and 77.9% on the DeepSWE v1.1 software-engineering benchmark [8]- solid numbers, but not the same as a verified Opus 5.5-beating claim for Carbon, which has no published scores at all.

Why the Best Version of Gemini 4 Might Never Ship

Google has not said whether Carbon becomes a future Argon update, a distinct Gemini 4 release, or stays an internal-only checkpoint that users never see [2]. That ambiguity matters because Argon itself was first released narrowly, through the Fairwind Program for trusted cyber-defenders, rather than to the public at large [8]. Wall Street has noticed Google's broader release pattern: JPMorgan analysts argue Google has 'struggled to translate ongoing product velocity into splashy releases that resonate with consumers, capture sustained mindshare, and meaningfully drive sentiment around AI leadership' [9], which is exactly the risk a leak like Carbon's runs into - impressive in private, invisible in the market until (or unless) it actually ships.

Excited Leakers, Skeptical Users: The Social Split Over Carbon

The reaction to Carbon splits cleanly by platform. On X and YouTube, AI-news accounts treated the leak with excited, speculative energy, lining Carbon up alongside other concurrent frontier-model rumors and framing it as a serious rival to Opus 5.5 and even GPT-6. One recurring detail from that side of the conversation: Google itself appears to have slipped Barium-B - the internal checkpoint that became the public Argon release - into the LMSYS Chatbot Arena disguised under a decoy model name to gauge real-world performance before any official reveal, a practice that shows Google was already testing checkpoints incognito well before the Carbon leak surfaced. Reddit's response ran in the opposite direction. The Gemini community there treated the leak as the latest entry in a tiring pattern of 'next checkpoint beats everything' stories, with one post bluntly titled to mock Google for teasing newer models before Argon itself was even broadly available to paying users. The most upvoted skeptical take argued Carbon is probably not a distinct model tier at all, just an internal build feeding into the existing Argon release - a far less exciting story than the Opus 5.5 comparisons circulating elsewhere. That gap between hype and skepticism is itself revealing: it shows an audience that has been burned before by leak cycles that never translate into a shipped product, sitting right alongside a louder crowd eager to anoint the next frontier model before Google has said a word.

Historical Context

2025-11-18
Google Antigravity, the agentic development platform tied to the internal Jetski name, was announced alongside Gemini 3, initially powered by Gemini 3.1 Pro and Gemini 3 Flash.
2026-09-23
DeepMind chief Koray Kavukcuoglu said Gemini 4 had moved into post-training and was being tested internally in Antigravity, with hopes to release an early post-training version before the end of 2026.
2026-09-30
Google announced Gemini 4 Argon, its new flagship model, initially releasing it only to trusted members of its Fairwind Program for cybersecurity use before wider rollout.
2026-10-09
Business Insider published an exclusive report, citing internal documents, screenshots, and conversations, revealing that Google employees were internally testing a newer Gemini 4 checkpoint called Carbon on Jetski.

Power Map

Key Players
Subject

Google is internally testing a new Gemini 4 checkpoint codenamed Carbon on its Jetski coding platform, with at least one employee describing its coding performance as comparable to Anthropic's Claude Opus 5.5 - though Google has confirmed no benchmarks, release date, or public availability for Carbon.

GO

Google / Google DeepMind

Developer of Gemini 4, Argon, and the internal Carbon checkpoint; operates the Jetski internal coding platform tied to Antigravity

BU

Business Insider

Original reporting outlet that broke the Carbon story via internal documents, screenshots, and employee conversations

AN

Anthropic

Maker of Claude Opus 5.5, the coding model Carbon is being compared against by Google employees

GO

Google employees (anonymous internal testers)

Primary source of the Carbon performance claims via internal messaging channels, not an official statement

JP

JPMorgan analysts

Wall Street analysts commenting skeptically on Google's ability to translate AI product velocity into market-moving releases

Fact Check

9 cited
  1. [1] Google Is Reportedly Testing a New Gemini 4 Checkpoint Called Carbon
  2. [2] Gemini 4 Carbon Leak
  3. [3] Google's Gemini 4 Carbon model is reportedly matching Anthropic's Opus 5.5 coding performance
  4. [4] Google Antigravity
  5. [5] Google Gemini 4 Post-Training Antigravity Release
  6. [6] Google's Gemini 4 Carbon checkpoint internal reactions
  7. [7] Google employees test new Gemini 4 model codenamed Carbon
  8. [8] Google launches Gemini 4 Argon, its most advanced AI model yet
  9. [9] Google Gemini 4 arrives as Wall Street shifts to personal agents

Source Articles

Top 5

THE SIGNAL.

Analysts

“Said Carbon's coding ability feels on par with Anthropic's Opus 5.5, but added that more testing is needed before giving a firm verdict.”

Unnamed Google employee (quoted by Business Insider)
Internal tester of Carbon

“Noted that earlier Argon versions reminded them more of Anthropic's older Opus 5 on some coding tasks, implying Carbon marks a real step up rather than a universally shared impression.”

Another unnamed Google employee
Internal tester, more cautious view

“Argued Google has struggled to turn frequent model releases into headline-grabbing launches that shift public and investor sentiment on AI leadership.”

JPMorgan analysts
Skeptical of Google's ability to convert AI progress into market impact

“Questioned whether Business Insider correctly understood what a model checkpoint actually is, reflecting broader skepticism about the technical rigor of the leak coverage.”

Anonymous online commenter (Digg thread)
Skeptical of reporting accuracy and terminology
The Crowd

“Business Insider is reporting that Google is internally testing a new Gemini 4 checkpoint named Carbon that matches Opus 5.5 in coding.”

@@AndrewCurran_1052

“Google hasn't even rolled out Gemini 4 Argon to the public yet, and its engineers are already testing a MUCH more powerful successor! According to Business Insider, @GoogleDeepMind has been testing several Gemini 4 variants internally, codenamed Argon, Barium, and Carbon.”

@@mark_k258

“Google Gemini 4 Carbon checkpoint is on par with Opus 5.5. we have not got argon yet lol, when they will launch carbon checkpoint we will have new sota from anthropic and openai. Please google launch your model, we are without a pro launch since 3.1, learn from xAI.”

@@chetaslua262

“Google announced Gemini 4 Carbon but Gemini 5 is even better!”

@u/OddBig010380
Broadcast
HUGE Gemini 4 Carbon LEAKS, Claude Fable 6, GPT-7 Bel, Google CRACKED RSI?! & More! AI NEWS

HUGE Gemini 4 Carbon LEAKS, Claude Fable 6, GPT-7 Bel, Google CRACKED RSI?! & More! AI NEWS

New Gemini 4 Pro Leak Shocks Everyone With Huge Upgrade

New Gemini 4 Pro Leak Shocks Everyone With Huge Upgrade

Gemini 4 Carbon, Google's Next Model? Fable 5.5 Next Week

Gemini 4 Carbon, Google's Next Model? Fable 5.5 Next Week

Google is internally testing a new Gemini 4 checkpoint codenamed Carbon on its Jetski coding platform, with at least one employee describing its coding performance as comparable to Anthropic's Claude Opus 5.5 - though Google has confirmed no benchmarks, release date, or public availability for Carbon. — AI News | Agentic Brew