In a recent LinkedIn post, Rand Fishkin delves into the nuanced relationship between specific publications and the way Large Language Models (LLMs) process and present information. Fishkin addresses a key challenge for content creators and SEO professionals: understanding which sources are most likely to influence LLM outputs, particularly for Retrieval-Augmented Generation (RAG) systems.
Fishkin highlights the limitations of LLMs themselves in answering this question directly. He states:
“Wait, you guys can tell me which publications (in my field) are likely to impact the responses LLMs give?”
He then clarifies the capabilities and limitations of these AI models in this context. According to Fishkin, while LLMs may struggle to definitively identify the most impactful publications, they can be instrumental in filtering a pre-existing list of audience-frequented sources.
Understanding LLM Crawling and Indexing
Rand Fishkin points out that the effectiveness of a publication in influencing LLM responses hinges on whether it is crawled and indexed by the models. This process is crucial for the information to be incorporated into the LLM’s training data or accessible for RAG systems.
As Fishkin explains:
“Yes. Yes, we can. This is because LLMs are not great at telling you the answer to this question. But they are good at filtering a list of publications your audience already reads/visits into those more and less likely to be crawled/indexed by training models AND come up highly in search (for RAG).”
This suggests a two-step approach for leveraging this insight. First, identify the publications your target audience engages with. Second, use the LLM’s filtering capabilities to determine which of those are most likely to be part of the LLM’s knowledge base or readily accessible for real-time information retrieval.
Strategic Publication Selection for SEO and RAG
Fishkin’s analysis implies a strategic imperative for businesses and content creators. To maximize the impact of their content within LLM-driven search and information synthesis, they must consider not only audience relevance but also the technical aspects of LLM data ingestion.
The Importance of Crawlability and Indexability
The core of Fishkin’s argument is that the ‘impact’ of a publication on LLM responses is directly tied to its visibility within the LLM’s operational framework. Publications that are regularly crawled and indexed are more likely to be part of the data set that LLMs draw from, whether for their foundational training or for real-time information retrieval in RAG applications.
According to Rand Fishkin, the goal is to identify sources that are not only authoritative and trusted by humans but also recognized and processed by AI. This dual requirement is essential for ensuring that content has the potential to surface and influence the answers provided by LLMs.
In essence, Fishkin is guiding professionals to think critically about the technical SEO of their content’s distribution channels, understanding that the digital ‘pipes’ through which information flows to LLMs are as important as the quality of the information itself.
📝 About This Content
This article is based on insights shared by Rand Fishkin on LinkedIn.
📅 Originally posted on May 5, 2026 | View original post on LinkedIn →