Gemini 4 Pro Allegedly Launched Anonymously, Benchmark Scores Surpass GPT-6 and Claude

nashnova research
今天发布阅读约 9 分钟

An anonymous model suspected to be Google's next flagship Gemini 4 Pro has surfaced on an AI leaderboard, outscoring GPT-6 Astra and Claude Fable 5.1 across multiple benchmarks — if confirmed, Google reclaims the performance crown after a half-year drought.

01

How was this "anonymous model" discovered?

A model tagged "gemini-3.8-flash" appeared on the Chatbot Arena leaderboard — but the real Gemini 3.8 Flash shipped back in early September.
This means → the same name showing up again is highly irregular. Multiple developers who tested it concluded it is likely an early checkpoint of Gemini 4 Pro — a snapshot from partway through training, running under a borrowed label.
In plain terms = Google is almost certainly stealth-testing a new model under an old name to avoid premature exposure.
02

How strong are the benchmark scores?

Agentic coding task DeepSWE v1.1: 88%, roughly 2 percentage points above GPT-6 Astra.
Real-world knowledge work GDPval-AA v2: 2,064 Elo — the only model among the three to reach that level.
Terminal coding Terminal-bench 2.1: 95.3%. Computer-use OSWorld-2.0: 86.8% — top-tier on both coding and OS operation.
None of these scores have been officially confirmed by Google; they come from independent developer testing only.
03

What do pricing and specs reveal?

Researchers who accessed the back end report: 10 million token input, 256,000 token output, persistent cross-session memory — the model remembers prior conversations — and built-in web access without an API call.
Input pricing: $2.25 per million tokens. Output: $11.25 per million tokens — reportedly below both Astra and Fable 5.1.
This means → if the numbers hold, Google is leading on performance *and* undercutting on price — competing on value, not just capability.
04

What is "recursive self-improvement"?

The prevailing market narrative: Gemini 4 Pro finished pre-training ahead of schedule because Google DeepMind achieved an RSI loop — Recursive Self-Improvement, where the model continuously refines its own training strategy during training, creating a positive feedback cycle.
DeepMind Chief Strategy Officer Jasjeet Sekhon publicly stated that RSI is now a key pillar of the investment thesis behind large-scale AI spending. Google also published a study called Dream-RSI, exploring how agents can improve their own search strategies while exploring.
In plain terms = ordinary AI gets smarter because humans feed it more data. RSI means the model teaches itself and accelerates its own training — which is why Gemini 4 Pro may have arrived earlier than expected.
05

What does this mean for Google's competitive standing?

Google's last flagship, Gemini 3.1 Pro, shipped in February — over six months ago. The planned June release, Gemini 3.5 Pro, was scrapped because it underperformed Flash.
This means → Google has been silent at the flagship tier for two product cycles. The market had begun to treat it as falling behind OpenAI and Anthropic.
If Gemini 4 Pro's benchmarks are ultimately verified, this becomes the inflection point where Google re-establishes flagship competitiveness — a narrative shift from "chaser" back to "leader."

市场有风险,内容仅供研究参考,不构成投资建议。

Gemini 4 Pro Allegedly Launched Anonymously, Benchmark Scores Surpass GPT-6 and Claude · nashnova