Google Develops Inference-Specific Chip, Potentially 10x More Efficient Than TPUs

0xBroomberg
Published 2026-07-20About 10 min read

Google is building a chip codenamed Frozen v2 that bakes parts of Gemini's decision logic directly into silicon, targeting 6–10x the tokens-per-watt of its latest TPU — a radical bet on locking chip to model to break through a worsening AI compute shortage.

01

What is Frozen v2, and how does it differ from a TPU?

The core idea: hard-wire part of Gemini's decision logic into the chip itself, cutting runtime compute steps and data movement. This means → the chip stops being a general-purpose calculator and becomes a purpose-built machine for one model family.
TPUs and Nvidia GPUs are general-purpose AI chips that can run many models. Frozen v2 can only run Gemini versions sharing the same base architecture it was designed around. In plain terms = a TPU is a kitchen that cooks anything; Frozen v2 is a machine that makes one dish — but makes it extremely fast.
The trade-off is flexibility: if Gemini's base architecture changes fundamentally, the chip could become obsolete.
02

6–10x more efficient — what does that number actually mean?

Google employees on the project estimate Frozen v2 can process 6 to 10 times more tokens per watt than the latest-generation TPU.
This means → under the same power and cooling budget, inference throughput multiplies several-fold — and for Google, that directly determines how many users it can serve.
Google plans to deploy the chip as early as 2028, but whether the efficiency gains hold up at production scale remains unproven.
03

Why is Google taking such an aggressive approach?

Google faces a severe AI compute shortage that has already sparked internal friction — Google Cloud has been forced to turn away some external customers.
Frozen v2's predecessor was led by DeepMind chief scientist Jeff Dean. The original Frozen design burned model weights directly into silicon, but was shelved because it locked the chip to a single Gemini version, making its useful life too short.
Frozen v2 pulls back from that extreme, retaining some flexibility — Google can still update the chip, including loading new model weights. In plain terms = it moved from "fully welded shut" to "semi-custom and tuneable."
04

How does Google plan to scale and position this chip?

Google has stated clearly that Frozen v2 production volumes will be far below TPU levels — it is a focused, specialized deployment, not a TPU replacement.
The company has not yet decided how much model information to hard-wire in. This means → the balance point between efficiency and flexibility is still open, and the final product could look quite different.
A Google spokesperson said "not every project makes it to production, but this rigorous exploration is core to our full-stack approach." This reflects a posture of exploration, not commitment.
05

Who else is chasing inference-specific chips?

Competition in AI inference chips is accelerating: startups SambaNova and d-Matrix, plus giants OpenAI and Microsoft, are all building dedicated inference silicon — each aiming to run inference cheaper than Nvidia GPUs.
Nvidia itself moved to close the gap, spending $20 billion last December to license technology from inference-chip startup Groq.
The closest conceptual peer to Frozen v2 is Canadian startup Taalas, which also hard-wires a specific AI model into silicon and has raised over $200 million from Quiet Capital, Fidelity, and others. This signals that "chip-locked-to-model" is not Google's idea alone — it is a thesis the industry is actively testing.

Content is for reference only, not financial advice.

Google Develops Inference-Specific Chip, Potentially 10x More Efficient Than TPUs · nashnova