In a recent LinkedIn post, Ann Smarty raises intriguing questions about how large language models (LLMs) interact with and understand major online platforms like Reddit, Wikipedia, and YouTube. Smarty, a recognized figure in the digital marketing and SEO space, shares her observations and concerns regarding the current state of AI’s comprehension of these influential content hubs.
She begins by highlighting the consistent strength of certain platforms in the AI landscape, noting:
“Reddit, Wikipedia, and Youtube win everywhere (except ChatGPT doesn’t like/understand Youtube that much; and AI Mode doesn’t like Wikipedia anymore, which is surprising as it still ranks, and it’s the foundation of Google’s knowledge overall… We all know how Google’s Knowledge Base came to be, right? Yeah, it’s scraped Wikipedia ๐ )”
AI’s Struggle with Visual and Foundational Content
Smarty points out a curious dichotomy where platforms like Reddit and YouTube, which are rich in user-generated content and video, seem to pose challenges for certain AI models. The observation about ChatGPT’s apparent difficulty with YouTube content suggests a potential limitation in processing or interpreting video-based information. More surprisingly, Smarty notes AI Mode’s current disinterest in Wikipedia, a platform she identifies as foundational to much of the internet’s knowledge base, including Google’s own knowledge graph. She alludes to the fact that Google’s knowledge base is largely built upon scraped Wikipedia data, making AI Mode’s dismissal of it particularly noteworthy.
LinkedIn’s Role in AI Training Data
The discussion then shifts to LinkedIn, a platform often perceived as a source of professional insights. However, Smarty expresses skepticism about the quality of its content when used for AI training.
“LinkedIn is trusted by many LLMs (provided it’s 80% of AI-generated slop, I can see it is likely hurting training data)”
According to Smarty, despite its professional veneer, a significant portion of LinkedIn’s content may be AI-generated or of low quality, potentially compromising the integrity of the training data used by LLMs. This raises concerns about the reliability of AI models that heavily rely on data scraped from such platforms.
Questioning Perplexity’s AI Classification
In her post, Ann Smarty also shares a moment of uncertainty regarding the classification of Perplexity AI.
“I am not sure if I’d call Perplexity an LLM…”
This statement suggests that Smarty perceives Perplexity AI’s functionality or underlying technology as potentially differing from what is typically defined as a large language model. Her observation invites further consideration into the evolving landscape of AI tools and their precise categorization.
Overall, Ann Smarty’s LinkedIn post serves as a critical commentary on the current capabilities and data dependencies of AI, urging a closer examination of the sources and quality of information these models are trained on.
📝 About This Content
This article is based on insights shared by Ann Smarty on LinkedIn.
📅 Originally posted on April 1, 2026 | View original post on LinkedIn โ