Key Concepts
- Agent Observability: The practice of monitoring and evaluating the performance, quality, and reasoning of AI agents in production.
- Non-deterministic Systems: AI agents that do not follow fixed code paths, making their behavior variable and harder to predict compared to traditional software.
- Traces and Spans: A "trace" represents a full workflow interaction, while a "span" is a single step within that interaction (e.g., a model call or tool usage).
- Human-in-the-loop (HITL): The integration of human feedback/annotation to grade agent performance and create ground-truth data for automated scoring.
- Tantivy: An open-source, Rust-based search engine library used for full-text indexing of large, unstructured trace data.
- Evals: The process of testing agent performance in batch using known inputs to optimize the system offline.
1. Traditional vs. Agent Observability
Phil Hetzel distinguishes between traditional observability and agent observability based on scope and system behavior:
- Traditional Observability: Focused on uptime and technical performance. It measures system health using metrics like latency, duration, and error rates (400/500 level errors). It is designed for deterministic applications with known control flows.
- Agent Observability: Focused on functional quality and reasoning. While it includes technical metrics (time to first token, total tokens), it must also evaluate qualitative aspects:
- Grounding: Was the response based on the provided context?
- Tool Usage: Did the agent reason correctly using available tools?
- Alignment: Does the response adhere to brand standards and system prompts?
2. Technical Challenges of Agent Traces
Agent observability presents unique data engineering hurdles:
- Data Volume and Structure: Agent traces are "nasty"—highly semi-structured and containing massive amounts of unstructured text. A single span can reach 20MB, and a full trace can exceed 1GB.
- Real-time Requirements: AI engineers require instantaneous visibility into agent interactions to maintain product-market fit, necessitating high-speed ingestion and processing.
- Read Patterns: Systems must support both real-time monitoring and complex analytical queries (e.g., searching for every trace containing a specific keyword like "Amazon").
3. Methodology: Building a Custom Database
To solve these challenges, BrainTrust developed a custom database architecture rather than relying on standard OLAP solutions like ClickHouse:
- Write-Ahead Log: Ensures immediate visibility of traces.
- Full-Text Indexing: Utilizes Tantivy to allow for efficient searching across unstructured text within traces.
- Unified Interface: All data is accessed via a SQL-like language, bridging the gap between technical observability and qualitative evaluation.
4. The Role of Human Annotation
A critical differentiator in agent observability is the inclusion of non-technical stakeholders (e.g., clinicians, lawyers, or subject matter experts).
- Process: Experts review traces and provide qualitative grades or justifications for an agent's performance.
- Application: These human-annotated justifications are used to train LLMs to create "scalable scoring functions," effectively automating the detection of failure modes that were previously identified by humans.
5. Future Directions: From Observability to Insight
The industry is moving toward "unknown unknowns" discovery. By running lightweight LLMs over incoming traces, platforms can perform:
- Embedding and Clustering: Grouping traces to identify common user intents.
- Sentiment Analysis: Understanding how users feel about agent interactions.
- Topic Modeling: Automatically surfacing recurring issues or themes in production data to shorten the iteration loop between identifying a problem and deploying a fix.
Synthesis
The transition from traditional to agent observability is driven by the shift from deterministic code to non-deterministic AI models. Effective agent observability requires a hybrid approach that combines technical performance monitoring with qualitative, human-in-the-loop evaluation. By treating observability and "evals" as two sides of the same coin—one real-time and one batch-processed—teams can create a closed-loop system that continuously improves agent quality based on real-world production data.
AI summaries can miss context or contain errors. Check important details against the original video.





