Google Wants to Embed Gemini in Chip, Has AI Spending Pressure Been Resolved?

TL;DR
· The Information reported that Google is exploring a Gemini-specific inference chip, codenamed Frozen v2, which may be deployed as early as 2028.
· The debate revolves around whether hard-coding the model structure can offset the flexibility loss and time risk with higher efficiency.
· Related Topics: Alphabet, Nvidia, in-house chip supply chain for cloud providers, AI energy-efficient infrastructure.
Google is exploring a server AI chip, informally known internally as Frozen v2, aiming to directly incorporate some elements of the Gemini model into hardware to enhance inference efficiency. Following this news, Alphabet's stock surged over 3% intraday on July 20, but the closing gains narrowed.
The key point of this news is not that Google is developing another in-house chip. Investors are more concerned about whether, after the increasing cost of AI services, Alphabet can reduce the electricity, computing power, and data center pressure per question not just by continuing to increase capital expenditure but by binding models and hardware.
The target given in the report is quite ambitious. Frozen v2 aims to increase the number of tokens per unit of power service by 6 to 10 times over the existing or latest reported TPU. A token can be roughly understood as the basic unit for AI processing of text. This figure is still a goal of an undisclosed project and not a production result officially confirmed by Google.
Stock Price Trading Imagines the Cost Curve
The market is driven by optimism, trading not short-term performance but long-term cost scenarios. The most sensitive issue for AI investors now is whether the capital expenditure of big model companies will continue to rise and whether cloud business gross margins will be continuously eroded by inference costs.
Google's pressure is not abstract. In the same report, the project aims to alleviate the AI computing power shortage, and Google Cloud had rejected some external contracts due to capacity issues. Whether these capacity constraints directly led to Frozen v2 or not, they all indicate a shift: computing power is no longer just about investment growth but is also starting to become a constraint on revenue realization.
When ordinary users query Gemini, it consumes inference computing power. The larger the model, the more users, and the longer the responses, the more chips, electricity, and cooling the data center needs. For investors, this will ultimately translate into capital expenditures, cloud services gross margins, and model service pricing.
The market implication of Frozen v2 is not just "Google immediately solving the hash rate shortage." It is more like the long-term technical roadmap Alphabet has laid out: if more tokens can be serviced per watt of electricity, Google has the opportunity to handle more AI requests under the same power and data center constraints.
Embedding Models into Circuits Saves on Transportation and Scheduling
The core of Frozen v2 is hardcoding. It can be understood as follows: a general-purpose chip is like a versatile kitchen that can cook any dish. Frozen v2 is more like directly writing part of Gemini's recipe into the stovetop, reducing the steps of looking up the recipe, moving tools, and temporarily adjusting processes each time.
AI inference involves not only computation but also a significant amount of data transportation and scheduling. During model runtime, chips need to constantly read parameters, arrange operation paths, and move data between different storage layers. These actions themselves consume power, time, and chip resources.
If certain model structures are stable enough, they can be turned into fixed circuits. The benefit is shorter paths, less scheduling, and lower power consumption. However, the trade-off is straightforward — once a chip is designed according to a certain model structure, it is not easy to adapt to various new models like GPUs or TPUs.
Therefore, the 6 to 10 times efficiency target needs to be deconstructed. It refers to the number of tokens that can be serviced per unit of power, which does not equate to a 6 to 10 times reduction in Google's future data center overall costs. Real chips, real workloads, and real deployment scales will all reshape the financial implications of this number.
The timeline also needs to be downgraded in understanding. Reportedly, Frozen v2 may not be deployed until 2028 at the earliest, with project details still being determined. Before then, its impact on Alphabet's valuation is more akin to a long-dated option rather than a cost improvement already reflected in the profit and loss statement.
Specialized Chips Challenge the General Narrative but Do Not Equal Replacement
What's worth noting about Frozen v2 is that it brings the contradiction of the AI chip roadmap to the forefront: it is difficult to maximize both generality and extreme efficiency simultaneously. In recent years, the advantage of GPUs and TPUs has been the ability to adapt to different models and workloads, especially during the rapid model iteration phase.
Google's explored direction is reportedly more radical. Since Gemini is its flagship model, can a more tailored inference chip be designed specifically for it? As AI inference transitions from demonstration to daily service, efficiency improvement is no longer just an engineering metric but a part of the business model.
This will pose a marginal challenge to NVIDIA's narrative but does not yet amount to a "NVIDIA crisis." NVIDIA's moat still lies in its general ecosystem, software toolchain, and training clusters. If Frozen v2 holds value, it is more within specific Gemini inference scenarios at Google, reducing marginal reliance on general-purpose compute.
It also does not mean that the TPU has been marginalized. By report caliber, Frozen v2 is a new branch outside of the TPU, more like an efficiency tool tailored for the Gemini family. The TPU will still undertake a broader set of training and inference tasks, serving different models and cloud customers.
A more accurate assessment is that when a model is significant enough, has a high enough call volume, and the architecture is stable enough, dedicated chips start to make sense economically. It is not a signal of the industry's comprehensive shift to dedicated chips, but rather top model companies beginning to calculate separately for high-frequency inference scenarios.
Gemini Stability Determines the Value of This Option
The biggest risk of Frozen v2 lies not only in whether the hardware team can create a more efficient circuit, but also in whether Gemini's underlying architecture can stabilize enough to justify being hardcoded into the chip. AI models are still iterating rapidly, and what is optimal today may not be the best solution two years from now.
If there are significant architectural changes to Gemini in the future, Frozen v2 may need to be redesigned. This would elongate the timeline and weaken the cost advantage brought by hardcoding. The deeper the specialization, the greater the efficiency potential, but the higher the cost of locking in the technological path.
What the market needs to validate is whether this path can progress from early project clues to harder engineering nodes. The official roadmap, chip tape-out progress, mass production timing, actual efficiency caliber, and deployment scale will all determine whether it becomes a new anchor on Alphabet's cost curve or a forward experiment whose valuation has already been factored in.
For Alphabet, the current significance of Frozen v2 is to show the market that the pressure from AI expenditures may not necessarily have to be addressed solely by continuing to increase capital expenditures. However, until 2028, it is more like an option. What can truly change the valuation is not just the code name itself, but whether this chip can translate Gemini's high-frequency calls into measurable cost reductions.
Recommended
Eight-Year Investment U-Turn: Why Did Ethereum Suddenly Abandon Poseidon?
Aug 16, 10:00
The Wall Street Journal: How is AI Trading Stealing the Limelight from Cryptocurrency?
Aug 15, 14:00
Tencent Still Has a Dream
Aug 15, 11:27
To Catch North Korean Hackers, They Set Up a Fake Project
Aug 15, 10:00
From Litigation to Settlement: Positive Signal Released by HTX's Negotiation with FCA
Aug 14, 19:32
11,742 Shipping Addresses Exposed Alongside Trezor Orders
Aug 14, 19:01