Google’s AI Chip Strategy: Moving Beyond NVIDIA’s Dominance, According to Linas Beliūnas

L

Linas Beliūnas

LinkedIn Author

Building a Safer Internet with AI 🤖 | Scouting for top startups to invest in 💸 | The only newsletter you need for Finance & Tech at 🔔linas.substack.com🔔 | Financial Technology | FinTech | Artificial Intelligence | VC

In a recent LinkedIn post, Linas Beliūnas highlights Google’s strategic moves in the artificial intelligence hardware sector, suggesting the tech giant is preparing to challenge NVIDIA’s stronghold, not by direct competition, but by fundamentally altering the landscape of AI infrastructure.

Beliūnas points out that Google is reportedly designing new AI chips, specifically for inference tasks, in partnership with Marvell. This move, as detailed by The Information, signifies a shift in focus from raw compute power to memory efficiency, which Beliūnas identifies as the next critical bottleneck in AI development.

“AI is not bottlenecked by compute anymore. It’s bottlenecked by memory. And whoever fixes that, wins the next phase big time.”

According to Beliūnas, the inference stage of AI processing is becoming increasingly dominant in terms of cost, accounting for an estimated 70-80% of total AI compute expenditure. He argues that while NVIDIA’s GPUs offer flexibility, their cost at hyperscale can be prohibitive. In contrast, custom-designed Application-Specific Integrated Circuits (ASICs) can offer significant cost and performance advantages for specific workloads.

The Inference War: A Long Game for Google

Linas Beliūnas frames the current AI development as a multi-stage conflict, with training being the initial battleground and inference being the more protracted and strategically crucial war. He emphasizes that success in the inference phase hinges on controlling the entire technology stack.

Owning the Full Stack

Beliūnas elaborates on this critical strategic advantage, stating:

“And inference rewards one thing: Owning the full stack. Models → compilers → chips → data centers. Google has been building that since 2015. Everyone else is still renting it.”

This integrated approach, according to Beliūnas, allows Google to optimize every component, from the software compilers to the physical data center infrastructure, for maximum efficiency in serving AI models at scale. This contrasts with companies that rely on third-party hardware and software, potentially leading to inefficiencies and higher costs.

Beyond NVIDIA: A Strategic Diversification

The article posted by Beliūnas suggests that Google’s objective is not necessarily to outperform NVIDIA in every metric but to change the terms of engagement. By developing its own inference-optimized TPUs and MPUs, Google aims to reduce its reliance on existing chip suppliers like Broadcom and gain greater control over its AI infrastructure.

As Beliūnas notes, the focus is shifting from a simple comparison of model performance to a broader consideration of infrastructure efficiency and profit margins.

“Most importantly, the AI race is no longer model vs. model. It’s infrastructure vs. margins. And Google is optimizing for both.”

He concludes by commending Google’s long-term vision, likening their strategic maneuvering to a game of chess, while suggesting that many competitors are still focused on more immediate, less strategic moves.

📝 About This Content

This article is based on insights shared by Linas Beliūnas on LinkedIn.

📅 Originally posted on April 19, 2026 | View original post on LinkedIn →