Synthetic Data vs. Synthetic Users: Teresa Torres Clarifies Key AI Distinctions

T

Teresa Torres

LinkedIn Author

Author, Speaker, Product Discovery Coach

In a recent LinkedIn post, product discovery expert Teresa Torres addresses a critical distinction in the realm of artificial intelligence development: the difference between synthetic data and synthetic users. Torres emphasizes that these are fundamentally different concepts, each serving a distinct purpose in the AI product lifecycle, and warns against conflating their applications.

Torres begins by explaining the role of synthetic data in testing AI models, particularly large language models (LLMs). She notes that when building AI products, especially those that interact via natural language, developers need to anticipate a wide array of user inputs. To prepare for this variability before real customer data is available, teams can generate synthetic data.

“The purpose of synthetic data is to uncover error cases that you need to fix before you release to real humans. The goal is not to use it predict what your customers want or need. Think about it as stress testing your AI app. You are evaluating how well does it perform across various input.”

This synthetic data, Torres clarifies, is generated by asking an LLM to create a diverse set of potential inputs. For instance, an AI app designed to build workout plans might receive varied user responses to a question about available workout time, ranging from specific durations like “30 minutes” to more ambiguous phrases such as “not long” or “the norm.” According to Torres, generating and testing against these varied inputs is crucial for stress-testing the AI’s robustness and identifying potential failure points.

Understanding the Limits of Synthetic Data

Torres strongly advocates for the use of synthetic data as a quality assurance tool. As she points out, its primary function is to ensure the AI application performs reliably across different scenarios. She states:

“synthetic data – great for understanding the quality of your AI app, not for understanding your customers”

This highlights that while synthetic data can simulate the *variety* of inputs an AI might encounter, it should not be relied upon to predict actual customer needs or behaviors. Its value lies in its ability to reveal bugs and performance issues within the AI model itself.

The Unreliability of Synthetic Users

In contrast, Torres delineates synthetic users as a separate and less reliable concept. Companies offering synthetic user services claim that LLMs can simulate customer behavior and predict needs, thereby augmenting or even replacing traditional product discovery processes. However, Torres expresses significant skepticism regarding these claims.

She argues that there is limited evidence to support the effectiveness of synthetic users for understanding customers. Furthermore, she cites studies indicating that synthetic users are not yet suitable for industry-level application.

“But this is not the same as synthetic users. Companies that offer synthetic users claim that an LLM can simulate your customers behavior and predict their needs. Their promise is that you can use their services to replace or augment your discovery. Not only is there little evidence that this works, there are several studies that show that synthetic users are nowhere ready for industry usage (more on that in a January blog post).”

Torres summarizes this crucial distinction by stating:

“synthetic users – sold as a way to understand your customers, but not yet reliable”

Her analysis serves as a vital reminder for AI developers and product managers to carefully consider the purpose and limitations of synthetic data and synthetic users, ensuring they are applied appropriately to avoid misinterpretations and ineffective product development strategies.

📝 About This Content

This article is based on insights shared by Teresa Torres on LinkedIn.

📅 Originally posted on December 16, 2025 | View original post on LinkedIn →