The maturity phases of running evals — Phil Hetzel, Braintrust

AI EngineerAbout 3 min readMay 29, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Agent Quality: The primary objective of evaluations (evals), ensuring agents perform reliably, safely, and cost-effectively.
  • Evals vs. Unit Tests: Evals focus on high-level failure modes rather than exhaustive code coverage.
  • LLM-as-a-Judge: Using an LLM to evaluate the output of another LLM or agent.
  • Observability: The practice of monitoring agent performance in production to maintain confidence.
  • Trace: A record of an agent’s execution, including inputs, tool calls, and outputs.
  • CRUD-based Tools: Tools that Create, Read, Update, or Delete data in external systems.
  • Flywheel Effect: The iterative process of capturing production traces, identifying failures, and using them to improve the agent in an offline environment.

1. The Importance of Evals

Phil Hetzel emphasizes that evals are essential for mitigating reputational, financial, and compliance risks. Unlike traditional software unit tests, which aim for exhaustive coverage, agent evals should be targeted at specific failure modes. Because the potential failure space for AI agents is infinite, developers should prioritize directional progress over perfect, 100% accurate test results.

2. Maturity Model for Agent Evals

Hetzel outlines a continuum of four stages for organizations building agentic systems:

  • Stage 1: Getting Started (Vibe Checking):
    • Methodology: Start by running 10–20 inputs through the agent.
    • Human Annotation: Use subject matter experts (SMEs) to provide a "thumbs up/down" and, crucially, a written justification.
    • Actionable Insight: Capturing these justifications is vital for later scaling the knowledge into automated systems.
  • Stage 2: Scaling Knowledge:
    • Methodology: Transition from manual human grading to automated grading using the justifications gathered in Stage 1.
    • Technique: Use LLM-as-a-Judge to replicate human logic.
    • Warning: Do not trust LLM judges blindly; they must be evaluated themselves to ensure they align with human standards.
  • Stage 3: Accounting for Complexity (Tooling & State):
    • Challenge: Agents interacting with external systems (CRUD operations) create complex state-management issues.
    • Solution: Use trace-based evaluation. Since traces can be large, they can encapsulate the system state at the time of execution.
    • Advanced Technique: Use versioned queries (e.g., querying a vector database at a specific timestamp) to recreate the exact environment the agent experienced during production.
  • Stage 4: Advanced Automation:
    • Methodology: Implementing topic modeling at scale to automatically uncover new failure modes in production data.
    • Integration: Utilizing CLI tools and CI/CD pipelines to automate the "flywheel" of capturing production data and rerunning it in offline evaluation environments.

3. Frameworks and Methodologies

  • The Evaluation Primitive: Every eval consists of three components:
    1. Task: The agent or prompt under test.
    2. Data Set: The input examples used to trigger the task.
    3. Scoring Function: The logic (code or LLM) used to judge the quality of the output.
  • The Flywheel: A continuous improvement loop:
    1. Capture production traces.
    2. Identify failures (via human or automated tools).
    3. Bring examples into an offline environment.
    4. Rerun evals to validate improvements.

4. Key Arguments and Perspectives

  • "Vibe Checking" is valid: While often criticized, initial manual assessment is a necessary starting point for understanding agent behavior.
  • Deterministic vs. Non-Deterministic: While deterministic graders (code-based) are reliable for things like token counts or tool-call frequency, LLM-as-a-Judge is necessary for subjective quality.
  • Eval the Eval: Hetzel argues that because LLM-as-a-Judge outputs are discrete, you can create a "ground truth" data set to evaluate the judge itself, ensuring it remains aligned with human expectations.

5. Synthesis and Conclusion

The transition from proof-of-concept to production-grade AI agents relies on moving from manual "vibe checks" to a robust, automated evaluation flywheel. By capturing production traces, documenting human justifications, and eventually automating those judgments through LLM-as-a-Judge and versioned state queries, teams can systematically reduce risk and improve agent quality. The ultimate goal is to treat evals not as static tests, but as a continuous process of rerunning production scenarios to guide development.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.