Turning Fails into Features: Zapier’s Hard-Won Eval Lessons — Rafal Willinski, Vitor Balocco, Zapier

AI EngineerAbout 5 min readJul 1, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Zapier Agents: Agentic alternative to Zapier for automating business processes.
  • Data Flywheel: The process of collecting user feedback, understanding usage patterns and failures, building new features, and improving the product, leading to more users and more data.
  • Instrumentation: Recording data from agent runs, including tool calls, errors, and pre/post-processing steps.
  • Explicit Feedback: Direct user feedback, such as thumbs up/down ratings.
  • Implicit Feedback: Indirect signals from user interactions, such as copying a model's response or rephrasing a question.
  • LLMOps: Software for understanding and managing agent runs, including tracing LM calls, database interactions, and tool calls.
  • Evals: Evaluations or tests used to assess the performance of AI agents.
  • Unit Test Evals: Simple assertions to check specific aspects of an agent's behavior, such as tool call parameters or keyword presence.
  • Trajectory Evals: Evaluating an agent's entire run, including all tool calls and generated artifacts.
  • LM as a Judge: Using a language model to grade or compare results from evals.
  • Rubrics-Based Scoring: Using an LLM to judge runs based on human-crafted rubrics that describe specific aspects to pay attention to.
  • AB Testing: Comparing different versions of an agent with a small portion of users to measure user satisfaction and key metrics.

Collecting Actionable Feedback

  • Instrumenting Code:
    • Trace completion calls, tool calls, errors, and pre/post-processing steps.
    • Record data in the same shape as it appears in runtime to easily convert it to an eval run.
    • Mock tool calls with side effects in evals.
  • Explicit User Feedback:
    • Ask for feedback in the right context, such as after an agent finishes running.
    • Example: Showing a feedback call to action after a test run resulted in a "nice bump" in feedback submissions.
  • Implicit User Feedback:
    • Mine user interaction for implicit signals.
    • Examples:
      • Turning on an agent after testing it.
      • Copying a model's response.
      • User expressing dissatisfaction in the conversation (e.g., "stop slinging around").
      • User rephrasing a question.
    • Use an LLM to detect and group frustrations.
  • Traditional User Metrics:
    • Track metrics relevant to the business.
    • Look at users who churned and analyze their last interactions.

Understanding Agent Runs

  • LLMOps Software:
    • Use LLMOps software to understand agent runs, which involve multiple LM calls, database interactions, and tool calls.
    • Consider building internal tooling for domain-specific understanding and easy conversion of failures into evals.
  • Analyzing Data at Scale:
    • Aggregate feedback, cluster interactions, and bucket failure modes.
    • Identify tools that fail most often and problematic interactions.
    • Use reasoning models to explain failures by providing trace output, input instructions, and other relevant information.

Building Evals

  • Eval Hierarchy:
    • Unit Test Evals (base): Predict the n+1 state from the current state.
    • Trajectory Evals (middle): Evaluate the entire agent run, including all tool calls and artifacts.
    • AB Testing (top): Compare different versions of an agent with real users.
  • Unit Test Evals:
    • Focus on simple assertions, such as checking tool call parameters or keyword presence.
    • Use for hill climbing specific failure modes.
    • Beware of overfitting to existing models.
  • Trajectory Evals:
    • Evaluate the entire agent run, including all tool calls and generated artifacts.
    • Do not mock the environment; instead, mirror the user's environment and create a synthetic copy.
    • Slower to run than unit test evals.
  • LM as a Judge:
    • Use an LLM to grade or compare results from evals.
    • Ensure the judge is judging correctly and avoid introducing biases.
  • Rubrics-Based Scoring:
    • Use an LLM to judge runs based on human-crafted rubrics that describe specific aspects to pay attention to.
    • Example: Did the agent react to an unexpected error from the calendar API and try again?
  • Eval Types Summary:
    • Use LM as a judge or rubrics-based evals for a high-level overview of system capabilities and benchmarking new models.
    • Use trajectory evals to capture multi-turn criteria.
    • Use unit test evals to debug specific failures.

AB Testing and User Satisfaction

  • Don't Obsess Over Metrics:
    • When a good metric becomes a target, it ceases to be a good target.
    • A 100% score on an eval data set may indicate that the data set is not interesting.
  • Data Set Division:
    • Divide the data set into a regressions data set (to avoid breaking existing use cases) and an aspirational data set (for extremely hard cases).
  • Ultimate Verification Method: AB Testing:
    • Route a small portion of traffic (e.g., 5%) to the new model or prompt.
    • Monitor feedback, activation, user retention, and other key metrics.
    • Make educated guesses based on real-world user data instead of optimizing for imaginary numbers in a lab setting.

Synthesis/Conclusion

Building good AI agents and platforms for non-technical users is challenging due to the non-deterministic nature of AI and user behavior. The key is to establish a data flywheel by collecting actionable feedback, understanding usage patterns, and continuously improving the product. This involves instrumenting code, mining both explicit and implicit user feedback, and using LLMOps software to analyze agent runs. Different types of evals, including unit test evals, trajectory evals, and LM as a judge, can be used to assess agent performance. However, the ultimate goal is user satisfaction, which can be best measured through AB testing and monitoring key metrics.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.