On July 20, The Information reported that Google is developing a server chip internally codenamed Frozen v2, with a single goal: running its Gemini models far more efficiently. Google declined to confirm or deny the report and left a boilerplate statement. The market responded first — Alphabet shares climbed roughly 3% on Monday morning after the news, just ahead of the company’s earnings report later that week. For a story built on anonymous sources, that is an unusually warm reception, and it reflects where investor attention sits in 2026: not on new model demos, but on the power bill behind them.
A Chip Built for Gemini
The point is not “another TPU” but a change in design assumptions. Today’s TPUs are general accelerators serving all of Google’s workloads; Frozen v2’s reported positioning is custom-built for Gemini, with the model’s characteristics baked into the hardware premise. The Next Web described the route as “a chip with Gemini baked into the silicon.” Google’s response was textbook: “Our teams are constantly researching and experimenting with new innovations to deliver maximum performance,” plus the caveat that “not every project moves into production” — acknowledging the project exists while reserving room to cancel it. The statement leans on Google’s “full stack approach,” which is the honest framing: the company sells the model, the chip, and the data center, so co-designing all three is the logical end state.
The Efficiency Target: Six to Ten Times
The reported numbers: Frozen v2 is expected in 2028 and could be six to ten times more efficient than Google’s existing AI chips, measured in tokens generated per unit of power. The choice of metric is itself a signal — inference, not training, is the next battleground. Training spends a fortune once; inference spends a slightly smaller fortune forever. As application layers and agent workloads push token consumption up by orders of magnitude, the electricity cost per token directly determines serving margins, and agents multiply the problem by making each user request worth dozens or hundreds of model calls.
The Custom Silicon Arms Race
Frozen v2 is not an outlier. OpenAI unveiled its first custom chip in June — Jalapeño, an inference processor built with Broadcom, notable for being aimed at inference on day one. Anthropic has reportedly been discussing a chipmaking partnership with Samsung. The shared motivation is straightforward: reduce dependence on NVIDIA and ease the compute shortage. NVIDIA remains the default supplier for the industry’s training fleets, and when both training and inference demand exceed supply, custom silicon stops being a cost-optimization option and becomes supply-chain insurance. Each lab also carries different constraints — OpenAI and Anthropic rent most of their capacity, while Google owns its stack — yet all three are converging on the same conclusion: own the silicon that your own workload actually needs.
Capex and the Market Reaction
The backdrop is just as striking: Alphabet plans to spend $180–190 billion on its AI buildout, and investor unease about that spending has not gone away. The efficiency story around Frozen v2 answers that unease directly — not by spending less, but by producing more tokens from the same power, which is the only sustainable version of the argument. The 3% stock bump suggests the market is buying the narrative for now, helped by the timing ahead of earnings. The real test arrives twice: first in the financials this week, then in 2028, when a six-to-ten-fold efficiency claim has to ship as silicon rather than a slide.
What It Means for Developers
Three points. First, a 2028 timeline means no reason to wait on near-term choices: TPU and existing GPU routes remain the mainline for the next two years, and roadmap placeholders rarely survive contact with production planning. Second, the endpoint of an inference-efficiency race is cheaper tokens — the long-run cost curve of Gemini-family APIs belongs in any pricing comparison, alongside the current prices everyone quotes. Third, if “custom silicon for one model family” becomes the norm, models and infrastructure bind more tightly together, and migrating between providers gets harder at exactly the moment vendors push platform-specific features. A multi-vendor strategy — or at least an abstraction layer over your model calls — becomes more valuable, not less.
Sources
- Google is working on a new AI chip designed to make Gemini more efficient — TechCrunch
- Google plans new chip to run Gemini models more efficiently, The Information reports — Reuters
- Google is building a chip with Gemini baked into the silicon — The Next Web
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
