Google reportedly developing a new AI chip to run Gemini models more efficiently
Google is reportedly building a new kind of chip, and it’s not just another upgrade to its familiar TPU lineup. According to a report from The Information, first surfaced on July 20, 2026, Alphabet is developing a server chip internally codenamed “Frozen v2,” designed for one specific job: running its Gemini models faster and cheaper than anything in Google’s current hardware fleet. Engineers reportedly project the chip could be six to ten times more efficient than Google’s existing AI chips, measured in tokens generated per unit of power.
Google hasn’t confirmed the project outright, but it hasn’t denied it either—and the timing, arriving just two days before Alphabet’s Q2 earnings revealed the company’s first-ever negative free cash flow quarter, says a lot about why this kind of efficiency push matters right now.
What Makes Frozen v2 Different From a Regular TPU
To understand why this chip is notable, it helps to understand what it isn’t. Google has built general-purpose TPUs (Tensor Processing Units) for over a decade—accelerators capable of running virtually any AI model loaded onto them. That flexibility comes at a cost: every time a TPU processes a request, it has to compute the model’s architectural requirements from scratch.
Frozen v2 reportedly takes a different approach entirely. Rather than staying general-purpose, it would hardwire part of Gemini’s actual architecture—the structural blueprint that determines how the model routes and processes information—directly into the chip’s circuitry. In machine learning terms, “freezing” means locking a component permanently in place rather than leaving it flexible. The practical effect, according to the reporting, is a chip that needs to perform fewer calculations and move less data through memory for every response it generates, since it isn’t recomputing Gemini’s structure on every single request.
Google offered a carefully hedged response when asked about the project, describing its teams as constantly researching new innovations for performance and efficiency, while noting that not every project moves into production—language that neither confirms nor denies Frozen v2 specifically.
The Trade-Off: Speed for Flexibility
This kind of specialization isn’t free. Because Frozen v2’s efficiency gains reportedly come from baking part of Gemini’s architecture directly into silicon, the chip would only remain useful as long as Google keeps that underlying architecture intact. If Gemini’s core structure changes significantly in a future version, the chip’s advantage could shrink or disappear.
That’s also why Frozen v2 is reportedly being treated as a specialized branch of Google’s chip portfolio rather than a replacement for its general-purpose TPUs, with production volumes expected to fall well short of TPU-level scale. And unlike TPUs, which Google rents out to Cloud customers, Frozen v2’s design is specific to Gemini, meaning it’s unlikely to ever be offered externally—this is reportedly an internal efficiency tool, not a new product line.
Why Google Needs This Right Now
The timing of this report isn’t a coincidence. Google is reportedly grappling with a real capacity crunch: around March 2026, the company told Meta it couldn’t supply the full volume of Gemini computing capacity Meta wanted to purchase, forcing Meta to ask employees to conserve their own AI processing usage. That kind of constraint—being unable to meet demand you’ve already generated—is exactly the problem a dramatically more efficient chip would help solve.
It also lines up with Alphabet’s broader financial picture. The company’s Q2 2026 earnings, released July 22, showed capital expenditures of $44.9 billion for the quarter, nearly double the prior year, pushing free cash flow negative for the first time in the company’s public history. A chip that delivers six to ten times more tokens per watt would meaningfully change that math over time, letting Google serve more Gemini demand without a proportional increase in new infrastructure spending.
There’s also a chip-market angle. Nvidia currently controls roughly 85% of the GPU market used for AI workloads, and its hardware, while dominant, was originally built for graphics rendering rather than large language models—meaning it carries overhead that purpose-built chips don’t. At the scale Google now operates, a genuine 6–10x efficiency gap over general-purpose hardware isn’t a rounding error; it’s potentially billions of dollars in infrastructure costs.
Part of a Broader Custom-Silicon Trend
Google building specialized AI hardware isn’t new—the company has developed TPUs since 2015, and its newest general-purpose entries, the TPU 8t and TPUi, are already optimized separately for training and inference workloads respectively. What’s different about Frozen v2 is the degree of specialization: rather than optimizing for a workload category, it’s reportedly optimized for one specific model family.
This also fits a broader industry pattern of major AI companies developing custom chips to reduce their dependence on Nvidia. OpenAI, for instance, unveiled its own first custom AI chip earlier this year. As AI infrastructure spending climbs across the industry, building model-specific silicon is emerging as one of the clearer paths to meaningfully improving the economics of running these systems at scale, rather than simply buying more general-purpose hardware.
How the Market Reacted
Investors responded quickly and positively. Alphabet shares climbed roughly 3% in trading on the Monday the report broke, ahead of the company’s Wednesday earnings release—a sign that markets viewed a more efficient, purpose-built chip as a credible answer to the capacity and cost pressures Alphabet has been facing amid its AI spending surge.
Important Caveats
A few things are worth keeping in perspective before treating Frozen v2 as a done deal:
- It’s based on anonymous sourcing. The Information’s report cites unnamed sources, and Google has not officially confirmed the project exists.
- The efficiency numbers are projections, not benchmarks. The six-to-ten-times figure reflects engineer estimates, not results from a working, tested chip.
- The timeline is distant and tentative. Deployment is reportedly targeted for as early as 2028, and even that date is described as a reported goal rather than a confirmed commitment.
- It’s currently exploratory. Frozen v2 is said to be an internal research project at this stage, and not every such project at Google ends up shipping.
Conclusion
Frozen v2 represents a notable shift in how Google may be thinking about AI infrastructure: instead of only building bigger, more general-purpose chips, it’s reportedly exploring hardware custom-built for a single model family, trading some flexibility for a potentially dramatic efficiency gain. Given Google’s mounting AI infrastructure costs, a real capacity crunch, and a newly negative free cash flow quarter, the strategic logic behind a chip like this is easy to understand—even if the project itself remains unconfirmed, years from deployment, and still very much in the research phase.