
Alphabet is developing a custom server chip, internally referred to as Frozen v2, that bakes portions of the Gemini model family directly into silicon in an effort to lower the power cost of serving AI responses. Engineers cited in the reporting estimate the design could deliver six to ten times the efficiency of Google’s current Tensor Processing Units (TPUs) when measured by tokens generated per watt, though the chip is not expected to ship until 2028.
What changes for SEO when inference gets cheaper
Cheaper inference rarely reaches the front end of a search engine in ways a site auditor can detect, but the trajectory matters for anyone planning content and infrastructure budgets. If a 6-10x efficiency gain lands around 2028, query-cost pressure on the search stack eases, which generally correlates with more generous real-time indexing, fresher SERP features, and faster response loops on AI-generated answers. For now, audit as usual: log response times, monitor crawl latency spikes, and watch for AI Overview volatility on your priority templates. The chip behind the curtain is irrelevant to on-page work, but the cost curve behind it shapes how aggressively Google can afford to expand AI surfaces in search.
How Frozen v2 differs from a standard AI accelerator
Frozen v2 follows a co-design philosophy: rather than treating the model as software that runs on general-purpose AI hardware, it embeds selected parts of Gemini into the silicon itself. The same approach shows up across the industry. OpenAI announced its first in-house inference chip, called Jalapeño, in June, and Anthropic has been reported as discussing a new partnership with Samsung. Each lab is trying to control more of the stack underneath its flagship model family so that every response costs less power, less time, and less silicon area.
The efficiency claim, in concrete terms
The benchmark in play is tokens generated per unit of power, a direct measure of how much useful output a chip produces for each watt it draws. A six to ten times improvement on that metric means the same rack of hardware could serve six to ten times as many model responses for the same energy bill, or the same workload at a fraction of the operating cost. Google did not confirm or deny the figures. A spokesperson said the company “constantly researches and experiments with new innovations” and emphasized that “not every project moves into production,” framing the work as part of a broader “full stack approach” where hardware and software are designed together.
Why Alphabet is pushing on silicon now
Two pressures are converging. Internally, serving Gemini at scale consumes a growing share of Alphabet’s infrastructure budget, and every efficiency gain on the inference path flows directly to the bottom line. Externally, the AI accelerator market has historically been dominated by Nvidia, and reducing that dependency is a strategic priority for every major lab. Custom chips also let Google tune the hardware tightly to the workloads that matter to Gemini specifically, including the long-context and multimodal paths that general-purpose GPUs handle less efficiently.
Capital spending and the investor reaction
Alphabet told the market earlier in the year that it plans to spend between $180 billion and $190 billion on capital expenditures to support its AI strategy. Shares climbed roughly 3% on the Monday after the Frozen v2 report surfaced, ahead of Alphabet’s earnings release later in the same week. A credible path to six to ten times efficiency on a future chip helps justify that level of outlay by promising a lower inference cost per query once the hardware is in production.
What this signals about the AI chip race
The competitive front line is shifting away from raw training throughput and toward inference specialization. Frontier performance is migrating from the data center floor into the silicon itself, with each lab designing accelerators around its own model family. For site owners and SEO practitioners, the practical takeaway is straightforward: AI-powered search surfaces are likely to keep expanding, the cost of generating those answers is on a downward trajectory, and auditing for AI Overview presence, structured data health, and crawl responsiveness remains the right call while the hardware catches up.
FAQ
What is Frozen v2?
Frozen v2 is the internal name for an AI server chip Alphabet is developing to make serving Gemini responses more efficient. The design hardwires selected parts of Gemini into the silicon rather than running the model purely as software on general-purpose AI hardware.
When is Frozen v2 expected to ship?
According to the original reporting, citing anonymous engineers, Frozen v2 is not expected to arrive until 2028, placing it in the long-range category rather than as an immediate upgrade for current Gemini users.
How much more efficient is Frozen v2 expected to be?
Engineers cited in the report expect Frozen v2 to deliver six to ten times the efficiency of Google’s existing TPUs, measured by tokens generated per unit of power. Google declined to confirm or deny the specific figures.
