Press "Enter" to skip to content

Google tests internal “Carbon” Gemini model as it readies Argon rollout

Google is poised to make its Gemini 4 model, publicly named Argon, available to customers, even as employees inside the company are already trialing a newer internal version dubbed Carbon. The move reflects the firm’s rapid iteration cycle on its next‑generation AI models.

Internal testing and codenames

According to internal documents and screenshots reviewed by Business Insider, Carbon was released to Google engineers through the company’s internal coding platform, Jetski, in the past few days. The model appears to be an updated checkpoint of Argon, though Google has not confirmed whether Carbon will be marketed under the Argon name or as a separate Gemini 4 variant.

Google’s internal naming scheme often differs from public branding. One internal memo identified a model called “Barium‑B” as the version that will be released publicly as Argon. The same memo listed a series of Gemini 4 checkpoints named Argon, Barium and Carbon, indicating that the company is testing multiple iterations before finalizing a public offering.

Employees have been discussing the new checkpoints in internal messaging channels. One engineer wrote, “Feels comparable to Opus 5.5,” while another added, “Carbon is really good!” A third comment referenced a “New Gemini pro next model,” a label the team uses for upcoming updates.

Performance claims and market pressure

When Argon was announced in late September, Google said the model delivered frontier performance on several coding and knowledge‑work benchmarks, including tasks in legal and finance domains. The company also highlighted Argon’s focus on defensive cybersecurity and announced that the model would first be offered to partners in its Fairwind Program, an initiative designed to test new models for cyber vulnerabilities before a broader rollout.

Early internal feedback on Argon was generally positive, but some staff noted that the model lagged behind Anthropic’s Claude Opus 5 on certain coding tasks. By contrast, the Carbon checkpoint has been likened by an employee to Anthropic’s Opus 5.5, the firm’s most capable model for long‑term agentic coding. The employee cautioned that further testing is needed to confirm the comparison.

Google’s internal testing underscores the company’s effort to keep pace in the fast‑moving race for AI coding agents. Anthropic and OpenAI have taken early leads among developers seeking models optimized for complex engineering work. Google’s recent announcement of a “universal” Gemini AI agent for the workplace signals its intent to compete not only on raw model performance but also on integrated, agentic assistants.

Google declined to comment on the internal testing of Carbon or the potential public release of additional Gemini 4 checkpoints.