Rand Fishkin Warns Marketers About Misinterpreting AI Citation Data

R

Rand Fishkin

LinkedIn Author

Cofounder of SparkToro, Alertmouse, & Snackbar Studio. Author of Lost & Founder. Feminist. I love underdogs, cooking, & helping people do better marketing

In a recent LinkedIn post, Rand Fishkin discusses his concerns that marketers are being misled by a common interpretation of “AI citation research.” Fishkin, a prominent figure in the marketing and SEO community, argues that the way citation data is being presented and understood in relation to AI-generated content may lead to flawed strategic decisions.

Fishkin highlights a specific trend that has caught his attention: the perceived decline in the influence of certain sources, like Reddit, on AI responses. He notes that some analyses show significant drops in citations from platforms such as Reddit in AI-generated answers, leading to assumptions about their diminishing impact.

“We see things like ‘Reddit dropped from 8% to 3% of citations in AI answers,’ and think that Reddit posts must therefore have less influence on the brands ChatGPT, Claude, or Google present in their responses… But that’s not true!”

The Correlation vs. Causation Fallacy in AI Data

A central theme in Fishkin’s analysis is the critical distinction between correlation and causation. He contends that while citation data might correlate with the brands or information that appear in AI answers, it does not necessarily imply a direct causal relationship. This nuance is crucial for marketers attempting to gauge the effectiveness of their content strategies in the age of AI.

According to Fishkin, the appearance of a citation does not automatically mean it was the sole or primary driver for an AI model including specific information. He elaborates on the complexities introduced by different AI architectures, such as Retrieval-Augmented Generation (RAG) versus base models, which can process and present information differently.

“Citations are correlated, sure, but not causal, to which brands appear in AI answers (and that’s further complicated by RAG vs. base model).”

Understanding the Nuances of AI Response Generation

Fishkin emphasizes that the underlying mechanisms of AI content generation are more complex than simple citation counts might suggest. The way an AI model retrieves, processes, and synthesizes information from its training data and real-time sources significantly impacts its output.

He points to the work of Lily Ray, who offers a crucial clarification regarding the RAG approach. Ray’s observation, as relayed by Fishkin, suggests that in specific RAG scenarios, particularly those involving very recent or limited information, the AI might indeed rely exclusively and causally on the provided citations.

“p.s. Lily Ray points out that in some cases of RAG where information is very recently updated and/or quite limited, the response will pull exclusively and causally from the citations.”

Actionable Insights for Marketers

Fishkin’s post serves as a cautionary note for marketers who might be oversimplifying the data emerging from AI citation research. He urges a deeper understanding of the methodologies and limitations behind these analyses.

As Rand Fishkin argues, marketers should avoid making definitive strategic shifts based solely on citation trends without considering the broader context of AI’s information retrieval and synthesis processes. The insights shared in his #5MinuteWhiteboard video aim to provide a clearer perspective on these complex issues, encouraging a more informed approach to content strategy in an evolving digital landscape.

📝 About This Content

This article is based on insights shared by Rand Fishkin on LinkedIn.

📅 Originally posted on March 5, 2026 | View original post on LinkedIn →