In a recent LinkedIn post, product discovery coach Teresa Torres discusses the critical, yet often misunderstood, practice of AI evaluations for product teams. Torres emphasizes that while AI evals have gained significant traction, many teams still lack a clear understanding of their purpose and implementation, often due to the technical nature of existing resources.
“AI evals (short for evaluations) are methods for measuring whether an AI product or workflow is performing well.”
Torres, who has previously described AI evals as a “new discovery habit,” aims to demystify this process with a recently published in-depth guide. She points out that the existing literature on AI evals is frequently geared towards engineers or lacks the specificity needed for product teams to effectively implement them.
Understanding the ‘Why’ Behind AI Evals
According to Teresa Torres, the primary function of AI evals is to instill confidence in product teams regarding the performance of their AI applications. She argues that these evaluations are not merely a technical exercise but a fundamental aspect of quality assurance and user protection.
“Evals give teams confidence that their AI applications are doing what they expect them to do. They help teams maintain quality and catch issues before they reach users.”
Torres elaborates that by establishing clear metrics for performance, teams can proactively identify and rectify potential issues. This proactive approach, as highlighted by Torres, is crucial in the rapidly evolving landscape of AI product development, where unexpected behaviors can have significant user impact.
AI Evals as a Core Discovery Habit
In her post, Teresa Torres draws a parallel between AI evals and other established product discovery practices, positioning them as an integral part of the feedback loop essential for iterative development.
The Feedback Loop Analogy
As Teresa Torres notes, the process of conducting AI evals mirrors other key discovery habits that product teams rely on.
“Similar to other discovery habits like interviewing and assumption testing, evals can act as a feedback loop to ensure we are on the right track.”
This comparison underscores Torres’s view that AI evals are not an isolated technical task but a strategic practice that informs the broader product development lifecycle. By integrating evals into regular discovery routines, teams can ensure their AI products align with user needs and business objectives, much like traditional user interviews or assumption testing guide other product decisions.
Making AI Evals Accessible
Teresa Torres’s initiative to create a practical, hands-on guide stems from her observation that many product teams struggle with the concept of AI evals. She states her intention was to make the process “practical, hands-on, and easy to follow.” This effort aims to bridge the gap between the technical requirements of AI evaluation and the practical needs of product managers and their teams, empowering them to adopt this crucial practice with greater confidence and clarity.
📝 About This Content
This article is based on insights shared by Teresa Torres on LinkedIn.
📅 Originally posted on September 4, 2026 | View original post on LinkedIn →