Google Building Frozen v2 AI Chip Designed Exclusively to Accelerate Gemini Workloads.
In an ambitious push to push the boundaries of custom silicon performance, Google is reportedly developing a new generation of internal AI server processors codenamed "Frozen v2" The specialized chip is engineered with a singular objective: to serve as dedicated hardware for running Google’s flagship Gemini model family, optimizing compute pipelines directly at the silicon level to match Gemini's unique architectural requirements.
By co-designing the hardware alongside the model's software parameters, Frozen v2 aims to dramatically reduce redundant mathematical computations and minimize data transfer overhead (memory-bandwidth bottlenecks) during high-throughput inference tasks.
In response to the reports, a Google spokesperson issued a measured statement that neither explicitly confirmed nor denied the project:
"Google continuously experiments with novel innovations to deliver maximum efficiency for our users and enterprise customers. While some experimental projects may never reach commercial deployment, they allow our research teams to establish foundational breakthroughs such as tailored hardware-software co-design that inform our future infrastructure roadmap."
According to internal reports, the Frozen architecture is designed to deliver a massive performance leap, processing 6 to 10 times more tokens per watt than Google's current Tensor Processing Units (TPUs) at equivalent power consumption levels. However, Frozen v2 is not intended to replace Google's versatile TPU fleet entirely; instead, it will function as an application-specific accelerator for specialized, high-density workloads.
Current internal timelines indicate that the Frozen chip series is scheduled for initial operational deployment by 2028.
The Google 'Frozen v2' Silicon Blueprint
The Codename: Frozen v2, a custom application-specific integrated circuit (ASIC).
The Primary Mission: Purpose-built to run Gemini model workloads with maximum energy and compute efficiency.
Performance Metrics: Projected to deliver 6x to 10x higher token processing efficiency per watt compared to current TPU generations.
Architectural Strategy: Hardware-software co-design to eliminate redundant data movement and memory latency.
Deployment Timeline: Target deployment scheduled for 2028 as a specialized accelerator alongside existing TPUs.
The trend of Hardware-Software Co-design: In the past, chips like GPUs or TPUs were designed to be flexible (general-purpose) to run a variety of AI models. However, Google's secretly developing a chip "specifically designed to run Gemini" reflects that as AI models become increasingly large, using general-purpose chips will face power consumption and heat constraints (Power Envelope Constraints). Removing unnecessary computational parts at the transistor level to ensure the chip works 100% in conjunction with Gemini is the only way to reduce server electricity costs.
The problem of Data Center Energy Consumption: Currently, the majority of the cost of running large-scale AI isn't in the chip itself, but in "electricity costs" and "heat management." The fact that Frozen v2 can process 6 to 10 times more messages/data (tokens) using the same amount of power is a game changer that will allow Google to provide Gemini to billions of people simultaneously without causing data center power failures.
Comparison with competitors... Every major player is secretly developing their own chips to reduce their reliance on NVIDIA (e.g., Microsoft Maia, Amazon Inferentia, Meta MTIA), but Google's move to create a dedicated chip for a single model, like the Gemini, demonstrates a vertical integration strategy (controlling everything from chip architecture to user interface) similar to the one Apple successfully implemented with its Apple Silicon chips.
Source: CNBC

Comments
Post a Comment