In a recent LinkedIn post, Teresa Torres explores the innovative ways incident.io is leveraging artificial intelligence to automate incident response, effectively creating an “AI SRE.” Torres, known for her work in product discovery and continuous delivery, shared insights from a discussion with Lawrence Jones, Founding Engineer at incident.io, and Ed Dean, Product Lead for AI.
The post delves into how incident.io is moving beyond traditional incident management, which focuses on coordinating teams and communication during outages, to developing AI that can actively diagnose and resolve issues. Torres highlights the core challenge and opportunity: teaching AI to think and act like a Site Reliability Engineer (SRE).
“Now, they’re building something new: an AI SRE that can actually help diagnose and respond to incidents.”
The Evolution of Incident Response with AI
Teresa Torres outlines how incident.io’s journey began with simpler AI applications, such as summarizing incidents, and has evolved into a sophisticated multi-agent system. This system is capable of forming hypotheses, testing them, and even drafting potential fixes, all orchestrated within the familiar environment of Slack.
Automating Debugging and Context Retrieval
A key focus of the discussion, as reported by Torres, is identifying which aspects of the debugging process can be safely automated. The team at incident.io has found success in combining retrieval methods, deterministic tagging, and re-ranking to quickly access relevant contextual information. This approach, according to Torres, often proves more effective than complex vector-based setups.
“AI’s biggest impact comes from compressing time—identifying causes minutes instead of hours.”
Torres emphasizes that AI’s primary value in this domain is its ability to significantly reduce the time it takes to identify the root cause of incidents. This acceleration is crucial in high-stakes environments where every second of downtime can be costly.
Building Trust and Confidence in AI Systems
The article shared by Torres also addresses the critical aspect of human trust in AI-driven workflows. The incident.io team is focused on how to balance human oversight with AI confidence, particularly within high-stakes incident management scenarios.
“Time Travel” Evals and Reasoning Transparency
To measure the effectiveness of their AI, incident.io employs post-incident “time travel” evaluations. As Torres explains, this method allows teams to score the AI’s performance retrospectively, once the actual sequence of events and the true cause are known.
“Building trust in AI isn’t just about precision—it’s about showing reasoning and uncertainty in ways humans understand.”
This approach underscores a broader principle: building trust in AI requires more than just accuracy. It involves clearly communicating the AI’s reasoning process and its level of certainty, enabling human operators to understand and effectively collaborate with the system.
The insights shared by Teresa Torres highlight a significant advancement in incident management, showcasing how AI can function as a collaborative, intelligent partner in critical operational workflows. The focus on automating diagnosis, providing context, and fostering trust positions AI as a powerful tool for SRE teams aiming to minimize downtime and enhance reliability.
📝 About This Content
This article is based on insights shared by Teresa Torres on LinkedIn.
📅 Originally posted on November 6, 2025 | View original post on LinkedIn →