In a recent LinkedIn post, Melissa Perri highlights an innovative approach to solving the data scarcity problem for AI product development, specifically referencing a solution implemented by Vanessa Lee and the team at Shopify for their product, Sidekick.
Addressing the AI ‘Chicken-and-Egg’ Problem
Perri, a recognized expert in product strategy, shared a compelling example of creative problem-solving in the AI space. The core challenge, often referred to as the ‘cold start problem,’ arises when an AI system requires data to function effectively, but that data cannot be generated until the system is already in use. Perri frames this dilemma, quoting Vanessa Lee from Shopify:
“We had the cold start problem… we had no data, we had no example conversations.”
This situation presents a significant hurdle for training AI assistants. As Perri explains, the team at Shopify devised a way to overcome this by manufacturing their own data, a strategy she found to be a particularly creative use of AI.
Manufacturing Data for AI Training
Melissa Perri details the ingenious method employed by the Shopify team. To bypass the lack of real-world interaction data, they developed a sophisticated ‘merchant simulator.’ This simulator involved several key steps:
- LLM-Generated Questions: Initially, large language models (LLMs) were used to create a vast array of potential questions that merchants might ask. These questions were varied across different industries and stages of business development.
- Simulated Merchant Interactions: Subsequently, another LLM was prompted to act as a specific type of merchant—in this case, a new user. This LLM then engaged in simulated conversations based on the generated questions.
- Manual Grading and Ground Truth: Product managers played a crucial role by manually grading these manufactured conversations. This process established the ‘ground truth,’ defining the quality standards necessary for training their core LLM Judge.
Perri emphasizes the significance of this manual grading, noting it was essential for establishing the benchmarks for quality.
Creating a Self-Improving Feedback Loop
The innovation didn’t stop at data generation. Perri outlines how the manufactured data and the ‘ground truth’ laid the foundation for a continuous improvement cycle once real users began interacting with Sidekick.
According to Perri, the LLM Judge was then deployed to continuously evaluate live conversations. This ongoing assessment created a powerful self-improving feedback loop. She states:
“Once real users started using Sidekick, this LLM Judge continuously evaluated live conversations, creating a self-improving feedback loop.”
Perri acknowledges that synthetic user testing can sometimes yield mixed results. However, she posits that this Shopify example demonstrates its effectiveness when executed with careful consideration and planning.
Broader Implications for AI Development
The insights shared by Melissa Perri offer valuable lessons for product teams navigating the complexities of AI development. The Shopify team’s approach underscores the importance of:
- Creative problem-solving in the face of data limitations.
- Leveraging LLMs not just for output, but for data generation and simulation.
- The indispensable role of human oversight (manual grading) in establishing quality for AI training.
- Designing systems with built-in feedback mechanisms for continuous improvement.
As Perri concludes her post by asking, “How are you solving data scarcity challenges in your AI products?”, she invites further discussion on this critical aspect of modern product development. This case study, as highlighted by Perri, provides a robust blueprint for tackling the common ‘cold start’ problem in AI, proving that even without initial user data, innovation can pave the way for effective AI solutions.
📝 About This Content
This article is based on insights shared by Melissa Perri on LinkedIn.
📅 Originally posted on December 15, 2025 | View original post on LinkedIn →