Google Internally Tests Stronger Gemini 4 New Version, Employees Say It Rivals Anthropic's Flagship
nashnova research
Google is preparing to publicly release its Gemini 4 "Argon" model, but employees are already testing a stronger internal version codenamed Carbon — with some saying its coding ability approaches Anthropic's best. The real power of the Gemini 4 family may not have been shown yet.
Argon, Barium, Carbon — how does the Gemini 4 family fit together?
Google internally named the Gemini 4 series after the periodic table: Argon → Barium → Carbon, representing three progressively stronger versions.
The "Argon" headed for public release is actually the internal build codenamed Barium-B. In plain terms = what the public gets as "version one" is already the second internal iteration.
Carbon is the latest and strongest internal build, but Google has not decided whether to release it as an Argon update or as a separate model within the Gemini 4 family.
An employee called Carbon "like Opus 5.5" — how significant is that?
A Google employee told *Business Insider* that Carbon "feels like Opus 5.5." Opus 5 is Anthropic's flagship model built for sustained, autonomous coding tasks; 5.5 implies a step above it.
This means → if the assessment holds, Carbon has reached the current ceiling of the coding-agent field — AI that can independently write, debug, and ship code.
The employee cautioned that "more testing is needed." Other internal comments were more impressionistic — "Carbon is really great!" and "the new version feels quite good" — with no systematic benchmark data made public.
How did early Argon perform, and why does Carbon matter so much?
When Google announced Gemini 4 last month, it highlighted Argon's frontier-level scores on coding and knowledge benchmarks spanning law, finance, and other domains, and emphasized its defensive cybersecurity capabilities.
But internal testing showed early Argon trailed Claude Opus 5 — Anthropic's previous-generation flagship — on some coding tasks. This reflects that Argon was not a clean sweep across every dimension.
That is exactly where Carbon's value lies: it may close Argon's coding gap. Yet Google does not guarantee it will release Carbon publicly. This means → the product the public eventually receives may not be the strongest version Google has in hand.
What does this mean for the AI coding race?
In the AI coding-agent race, Anthropic and OpenAI have built a clear lead among developers. Google is the challenger.
The Gemini 4 series will first open to "Fairwind Program" partners for cybersecurity vulnerability testing before a broader public release.
Put simply = Google may already possess a coding model near the top of the industry, but closing the gap depends on two things: whether Carbon ships publicly, and whether real-world performance after launch lives up to internal praise.
市场有风险,内容仅供研究参考,不构成投资建议。
