Google Frozen v2 chip
Published on
5 min read

Google Set to Redefine AI Inference Efficiency With New Frozen v2 Chip

In Focus

  • Frozen v2 chip is designed to run Google’s Gemini AI models
  • The new chip represents a shift to single-generation efficiency
  • Google has not confirmed the project officially yet

Google is reportedly developing a server chip that can embed Gemini’s structural blueprint directly into the circuitry under a project dubbed Frozen v2. The project could become one of the largest single-generation efficiency achievements in AI inference hardware. Google’s Frozen v2 chip represents a shift from general-purpose flexibility to single-generation efficiency.

Join thousands of readers who receive the latest software reviews, expert comparisons, and industry news delivered straight to their inbox. Subscribe Now

What is Google’s Frozen v2 AI Chip?

Frozen v2 is a next-generation AI chip designed specifically to run Google’s Gemini AI models more efficiently. Unlike conventional AI chips that are designed to support different large language models, the Frozen v2 chip will reportedly embed parts of Gemini’s architecture directly into the chip’s silicon.

This is expected to make the hardware more specialized for Google’s AI workloads. Engineers at the tech giant claim the new chip could deliver up to ten times the token output per watt compared to the company’s newest Tensor Processing Units. If successful, Google’s Gemini AI chip could lower the cost of running large language models at scale significantly.

Google, which introduced specialized AI chips for AI training and inference, has not confirmed the project officially yet. However, the tech giant is reportedly looking to deploy the new Google custom AI chip in 2028.

Why is Google Shifting to Single-Generation Efficiency?

The general-purpose flexibility approach allows AI chips to handle different tasks without making them efficient in any of them. Data center chips, including those manufactured by Nvidia and Amazon, are designed to support any AI model.

But this flexibility comes with hidden costs. When these chips run AI models, they have to continuously make runtime decisions. Additionally, the processors must draw model data from external storage, which consumes more energy than arithmetic operations.

Since large language models contain billions of parameters, they exceed the capacity of a chip’s on-board cache, which means constant movement of data between the chip and the external memory. This problem can only be fixed by minimizing data movement, rather than faster hardware or additional chips.

Google’s AI inference chip fixes these problems by embedding Gemini’s architecture directly into the chip. This eliminates the need for runtime scheduling while combining multiple operations into a single hardware process.

What Makes Frozen v2 Critical for Google?

Google considers the Frozen v2 chip as critical in addressing its compute constraints. Currently, compute shortage is costing the tech giant a significant portion of its revenue.

We are compute constrained in the near term. Our cloud revenue would have been higher if we were able to meet the demand,” Google CEO Sundar Pichai noted as cited by Tech Times.

Google Cloud’s backlog has surged to approximately $462 billion after nearly doubling in one quarter. In January 2026, the Gemini API alone handled 85 billion monthly requests, representing a 142% increase within ten months.

Linda Hadley
Scroll to Top