Google’s Next AI Advantage: Specialized Chips for Cheaper Inference, According to Linas Beliūnas

L

Linas Beliūnas

LinkedIn Author

🔔linas.substack.com🔔 Daily Intelligence on Finance & AI | Scouting FinTech & AI Startups 🦄

In a recent LinkedIn post, Linas Beliūnas discusses a potentially significant shift in Google’s artificial intelligence strategy, focusing on the development of specialized AI chips. Beliūnas highlights Google’s reported work on a new chip, dubbed “Frozen v2,” which aims to dramatically improve the efficiency of running its Gemini AI models.

Beliūnas explains that this approach deviates from creating general-purpose AI chips. Instead, Google is reportedly exploring hardwiring specific aspects of Gemini’s architecture directly into silicon. This custom integration, as Beliūnas notes, is intended to minimize runtime decisions and data movement, thereby increasing the number of tokens processed per watt of energy consumed.

“Instead of building a general-purpose chip that can run many AI models, Google wants to hardwire parts of Gemini’s architecture directly into silicon.”

The Drive for Inference Efficiency

The core of Beliūnas’s analysis centers on the immense cost of AI inference at Google’s scale. While training a large AI model is a significant one-time expense, serving those models across multiple products like Search, Workspace, and YouTube incurs continuous costs. Beliūnas points out that for a company like Google, inference is not just an operational cost but also a critical factor for profit margins, capacity planning, and energy consumption.

According to Beliūnas, the anticipated efficiency gains from Frozen v2, potentially 6–10 times better than current TPUs by 2028, could offer substantial strategic advantages. These include enabling Google to serve more queries without a proportional increase in data center infrastructure, reducing the cost associated with each Gemini response, and potentially lessening reliance on external chip providers like Nvidia for inference tasks.

Strategic Implications for Google Cloud

Beliūnas further speculates on how these advancements could impact Google Cloud. The increased efficiency could allow Google to offer AI services that are either more cost-effective or more powerful to its cloud customers.

“And at Google’s scale, inference becomes a margin, capacity, and energy problem.”

The Trade-off: Flexibility vs. Specialization

However, Beliūnas also addresses the inherent trade-offs in this specialized hardware approach. A chip optimized for the current Gemini architecture might become less relevant if the AI model undergoes significant architectural changes before its projected 2028 release. This risk, Beliūnas suggests, underscores Google’s broader strategic bet.

Linas Beliūnas argues that Google’s integrated control over its AI models, software, custom chips, data centers, and distribution channels presents a formidable competitive advantage. While competitors might achieve parity in AI model intelligence, matching Google’s efficiency in serving those models at a global scale could prove significantly more challenging.

“But matching Google’s cost per answer could be much harder.”

The Future AI Moat

In conclusion, Beliūnas posits that the next significant competitive advantage in the AI race may not solely lie in developing the most intelligent Large Language Models (LLMs). Instead, he suggests the true moat could be the ability to deliver intelligence at a planetary scale at the lowest possible cost.

“And that’s exactly why the next AI moat may not be the best LLM. It may be the cheapest intelligence to serve at planetary scale.”

Beliūnas’s analysis provides a compelling perspective on how hardware specialization could become a key differentiator in the ongoing AI landscape, particularly for companies operating at the scale of Google.

📝 About This Content

This article is based on insights shared by Linas Beliūnas on LinkedIn.

📅 Originally posted on July 20, 2026 | View original post on LinkedIn →