ChatGPT Citation Patterns: Ahrefs Study Reveals Surprising Truths, Reports Ann Smarty

A

Ann Smarty

LinkedIn Author

SEO for 20+ years, Reddit Marketing for 15+ years, AEO/GEO 💪 Co-Founder of Smarty Marketing

In a recent LinkedIn post, Ann Smarty dives into the intriguing findings of a new Ahrefs study concerning how ChatGPT cites its sources, particularly highlighting a significant discrepancy in how the AI attributes information derived from various online platforms.

Smarty frames the issue by referencing recent discussions about AI-generated content, noting the New York Times’ report on the potential for ungrounded citations in Google’s Gemini. She then introduces the Ahrefs study, which sheds light on the complex process of how ChatGPT selects and cites URLs to formulate its answers.

“Although ChatGPT crawls dozens of pages to answer a single query, it only ends up citing ~50% of them.”

This statistic alone underscores the challenge for content creators aiming to have their work recognized by AI models. Ann Smarty emphasizes that the study indicates a lack of “magic bullets” for ensuring citation, suggesting that relevance and specific alignment with the AI’s internal processes are key, yet still not guarantees.

The Reddit Conundrum in AI Citations

A particularly striking finding discussed by Ann Smarty revolves around the extensive use of Reddit by ChatGPT. According to the Ahrefs study, while Reddit is heavily utilized for understanding topics, gauging consensus, and building context, it receives remarkably little credit.

“ChatGPT is using Reddit extensively to understand topics, gauge consensus, and build context—but it almost never gives Reddit the credit (67.8% of all non-cited URLs come from Reddit). I wonder what Reddit thinks of that…”

Smarty uses this point to question the implications for platforms like Reddit, which serve as a significant data source for AI training without commensurate attribution. This raises broader questions about fair use and recognition in the evolving AI landscape.

Key Factors Influencing ChatGPT Citations

Ann Smarty breaks down the core factors identified by the Ahrefs study that influence whether a URL gets cited by ChatGPT. She highlights that title relevance to “fanout queries” is a crucial element.

As Ann Smarty explains:

“Ultimately, if your URL and title don’t semantically align with the AI’s internal fanout queries, you’re less likely to get cited.”

This suggests that content creators need to think not only about the content itself but also about how its title and semantic structure might align with the underlying questions an AI is trying to answer. Furthermore, the study presented an interesting paradox regarding content freshness.

Freshness vs. Age in AI Citation

Counterintuitively, while AI models are often associated with processing the latest information, the Ahrefs study, as reported by Smarty, indicates that ChatGPT tends to cite older content more frequently than very fresh content. This could imply that established, authoritative content, even if not brand new, holds a significant place in the AI’s citation patterns.

Ann Smarty concludes by reiterating that the pages most likely to be cited are those whose titles and content closely match the questions ChatGPT is internally processing. The retrieval channel through which the content is accessed also plays a role.

The insights shared by Ann Smarty provide valuable context for content creators, SEO professionals, and anyone interested in the mechanics of AI information processing, underscoring the complexity and nuances involved in how AI models attribute the vast amount of data they utilize.

📝 About This Content

This article is based on insights shared by Ann Smarty on LinkedIn.

📅 Originally posted on April 21, 2026 | View original post on LinkedIn →