The Web’s Bargaining Power: Hiten Shah on AI Training Data and Cloudflare’s Role

H

Hiten Shah

LinkedIn Author

CEO of Crazy Egg (est. 2005)

In a recent LinkedIn post, Hiten Shah discusses the emerging conflict around AI training data, highlighting how the internet is beginning to push back against unfettered access. Shah frames this as a pivotal moment where the value of high-quality human content for AI model improvement has become undeniable, leading to increased pressure on access.

As Hiten Shah notes:

“The real AI conflict shows up in Cloudflare’s logs. Four hundred sixteen billion crawler hits were stopped in a few months, and that single number tells you how quickly the web is starting to push back.”

Shah argues that publishers, who for years treated crawler requests as background noise, are now realizing the significant economic value these requests carry. This value stems from the fact that high-quality human-generated content is crucial for the rapid improvement of frontier AI models. The moment this value became apparent, a quiet but persistent pressure regarding access to this content began to mount.

The Complicated Position of Creators in the AI Era

The strategic decision by Google to fuse its search and AI crawlers has placed content creators in a particularly complex situation, according to Shah. Publishers face a dilemma: maintaining visibility in search results means their work inevitably flows into AI training datasets, while blocking these crawlers risks a decline in search visibility.

Hiten Shah points out the implications of this decision, stating:

“Google fused its search and AI crawlers, which puts every creator in a complicated position. Stay visible in search and your work flows into training. Block the crawler and your visibility starts to fade.”

This tension, as evidenced by Cloudflare’s data, is creating significant friction across the entire web, Shah observes. The article he references highlights that Google possesses reach into several times more of the internet than any other AI player. This extensive reach provides a structural advantage that grows as AI models are trained on information accessible primarily through Google’s platforms.

Cloudflare’s Emerging Role

Shah identifies Cloudflare as a unique player in this landscape, possessing the internet placement necessary to potentially disrupt this advantage. By offering infrastructure that can interrupt the flow of data, Cloudflare can provide publishers with a more genuine choice regarding how their content is used for AI training.

According to Hiten Shah:

“Cloudflare is the rare company with enough placement on the internet to interrupt that advantage and give publishers a real choice.”

A New Era of Negotiation on the Open Web

This situation signals a fundamental shift, Hiten Shah argues, ushering in a new era for the internet. The open web is evolving to a point where it can actively bargain for the value of its content. Infrastructure companies, like Cloudflare, are developing the capabilities to enforce these bargains. Consequently, AI companies are beginning to understand that guaranteed access to vast amounts of data is no longer a given.

Shah concludes by emphasizing the significance of this moment:

“This story matters because it signals a new era. The open web is learning how to bargain. Infra companies are learning how to enforce those bargains. AI companies are learning that access is no longer guaranteed.”

The WIRED article mentioned in Shah’s post, he suggests, captures this transitional phase before new industry rules solidify. It offers valuable insight into the shifting leverage points and underscores why the future trajectory of AI development is increasingly dependent on behind-the-scenes negotiations that are often invisible to the public.

📝 About This Content

This article is based on insights shared by Hiten Shah on LinkedIn.

📅 Originally posted on December 5, 2025 | View original post on LinkedIn →