Google is preparing a specialized AI chip for Gemini: faster speeds and lower consumption in data centers
Alphabet is reportedly working on a new generation of AI accelerators dedicated to its Gemini models, with a fundamentally different approach compared to the general-purpose chips currently used in data centers. The project, internally codenamed Frozen v2, aims to incorporate part of the model architecture directly into the silicon, transforming the chip into a component specifically designed to execute Gemini in the most efficient way possible.
This information was reported by The Information. Google has not officially confirmed the existence of the project nor provided technical details, but stated that its teams are constantly working on new hardware and software solutions to improve performance and efficiency. The company also emphasized that not all experimental projects necessarily reach commercial production.
The concept behind Frozen v2 represents an evolution over traditional AI accelerators. Currently, most AI chips keep the models loaded in memory and use programmable circuits to execute different neural networks. This ensures great flexibility, but also incurs costs in terms of energy consumption and processing times due to the ongoing transfer of data between memory and computing units.
With Frozen v2, Google intends to fix part of the structure of Gemini directly in the hardware. The model could continue receiving updates through new weights, but its basic architecture would remain tied to the chip's design. Hence the name 'Frozen', or 'frozen'.
According to available information, Google is evaluating which portion of the model to physically integrate into the silicon. The goal would be to achieve significant inference acceleration, with an estimated efficiency between 6 and 10 times higher compared to the company's current custom AI chips, measured in the number of tokens processed per unit of energy consumed.
Frozen v2 is not expected to replace the internally developed TPU (Tensor Processing Unit) but to complement them as a more specialized line of accelerators aimed at very specific workloads. Such a strategy would enable the company to optimize the execution costs of Gemini models in its data centers, especially considering the growing demand for AI-based services.
The push towards proprietary chips also stems from difficulties related to the availability of computational capacity. The expansion of generative AI has greatly increased the demand for accelerators, creating supply constraints and high costs. According to reports, the pressure on computational resources has contributed to internal tensions within Google and even led Google Cloud to refuse some requests from new customers.
Energy efficiency has indeed become one of the central elements of competition in the AI sector. Managing large-scale models globally requires extremely expensive infrastructure, and even small improvements in consumption per single operation can translate into significant savings across thousands of servers.
A dedicated chip for Gemini could also reduce latency, an important aspect for real-time applications such as voice assistants, conversational tools, and AI agents. Reducing the steps necessary to process a request could allow for quicker responses compared to a system designed to support a wide variety of models.
Google's strategy follows a broader trend in the tech industry. Major AI companies are seeking to develop proprietary hardware to reduce dependence on NVIDIA, which has solidified a dominant position in the AI accelerator market in recent years. Other players are also pursuing a similar path. OpenAI recently announced its first processor dedicated to inference, while Anthropic has reportedly initiated discussions for a chip collaboration with Samsung.
However, the approach presents an obvious limitation: rigidity. AI evolves rapidly, and a chip designed around the current architecture of Gemini may become less competitive over the years. Google would have therefore chosen a compromise, keeping the weights of the model updatable while leaving the main structure fixed.
According to reports, production of Frozen v2 is not expected to commence before 2028, and initially, the project may be smaller than the TPU family. The company is thus also utilizing the chip as an experimental platform.