Key Concepts
System design, software development, troubleshooting, production incidents, AI software engineering tools, AI ops, LLMs, agentic systems, causal machine learning, reasoning models, swarms of agents, autonomous troubleshooting, mean time to resolution (MTTR), observability data.
The Evolving Landscape of Software Engineering
The speaker identifies three core areas of software engineering: system design, software development, and troubleshooting production incidents. AI tools are rapidly automating software development, potentially shifting the focus to system design. However, the speaker argues that troubleshooting is becoming increasingly complex due to:
- Reduced human context: AI-generated code makes it harder for engineers to understand the system's inner workings.
- Increased system complexity: Systems are becoming more complex, exceeding human comprehension.
This could lead to engineers spending most of their time on QA and on-call duties, a "grim reality" if troubleshooting isn't addressed.
The Current State of Troubleshooting: Dashboard Dumpster Diving
The speaker describes the current troubleshooting workflow as "dashboard dumpster diving." Engineers sift through thousands of dashboards from tools like Grafana, Datadog, Splunk, Elastic, and Sentry to find clues about production incidents. This is followed by staring at the code base, trying to connect the issue to a recent change (pull request, configuration change). This often involves bringing in multiple teams, leading to large incident channels and prolonged resolution times.
Limitations of Existing AI Approaches to Troubleshooting
The speaker critiques three existing AI approaches to troubleshooting:
- AI ops (Traditional Machine Learning): Statistical anomaly detection generates too many false positives due to the complexity and dynamic nature of systems. "More noise than signal."
- LLMs (Large Language Models): While useful for explaining individual logs, LLMs can't handle the massive scale of production data (terabytes, trillions of logs). Context windows are insufficient, and LLMs lack a good understanding of numerical data.
- Agentic Systems (e.g., React-style agents): These agents rely on runbooks, which are often outdated. A broad search using simple tools can be too slow (days to run), while incidents need to be resolved in minutes.
Traversal's Approach: Combining Statistics, Semantics, and Agentic Control Flow
Traversal aims to achieve "out of sample autonomous troubleshooting" by combining three key elements:
- Statistics (Causal Machine Learning): Identifies cause-and-effect relationships from data, distinguishing root causes from correlated failures. This addresses the "correlation isn't causation" problem.
- Technical Term: Causal Machine Learning - Using machine learning techniques to infer causal relationships from observational data.
- Semantics (Reasoning Models): Leverages the semantic understanding of LLMs to interpret log fields, metadata, and code.
- Agentic Control Flow (Swarms of Agents): Employs thousands of parallel agentic tool calls to exhaustively search telemetry data efficiently.
This combination allows Traversal to find promising leads and connect them to specific changes in the system.
Case Study: Digital Ocean
Matt, the first employee at Traversal, discusses a case study with Digital Ocean, a cloud provider. Before Traversal, Digital Ocean engineers faced a "frantic search" through millions of metrics and billions of logs to resolve incidents.
- Example: An incident message might indicate a "potential compromise of some host."
Traversal's AI now investigates incidents automatically, orchestrating a "swarm of expert AI sRes" to sift through petabytes of observability data.
- Result: Digital Ocean has seen a 40% reduction in mean time to resolution (MTTR).
Traversal provides engineers with:
- Identified issues (e.g., a deployment causing a cascade of issues).
- Relevant observability data.
- Confidence levels for root cause candidates.
- Explanations of its reasoning.
- An AI-generated impact map.
- The ability to ask follow-up questions.
Broader Applications and Team
Traversal works with a heterogeneous group of enterprise environments, connecting to various observability tools and processing trillions of logs. The speaker emphasizes that this is both an AI agents problem and an AI infrastructure problem. The principles of exhaustive search and swarms of agents can be applied to other domains like network observability and cybersecurity.
The Traversal team consists of AI researchers, dev tools experts, AI product engineers, and quant finance traders. The speaker highlights the team's collaborative spirit and shared goal of improving the lives of engineers.
Synthesis/Conclusion
The presentation argues that current approaches to troubleshooting production incidents are inadequate and that AI-powered solutions are needed to address the increasing complexity of modern software systems. Traversal's approach, which combines causal machine learning, semantic reasoning, and swarms of agents, offers a promising solution for automating incident resolution and reducing the burden on engineers. The Digital Ocean case study demonstrates the potential of this approach to significantly improve MTTR and enhance system resilience. The speakers emphasize the importance of both AI algorithms and the underlying infrastructure to handle the massive scale of observability data.
AI summaries can miss context or contain errors. Check important details against the original video.





