Break It 'Til You Make It: Building the Self-Improving Stack for AI Agents - Aparna Dhinakaran

AI EngineerAbout 3 min readJun 11, 2025Watch original
THE SUMMARYAI-generated

Agent Evaluation at Arise: A Deep Dive

Key Concepts:

  • Agent Evaluation
  • Tool Calling
  • Trajectory Evaluation
  • Multi-Turn Conversation Evaluation
  • Self-Improving Agents
  • Eval Prompt Iteration
  • Orchestration Worker Pattern
  • Traces
  • Bottleneck Identification

1. Introduction: The Challenges of Building and Evaluating Agents

  • Building agents is complex, involving constant iteration on prompts, models, and tool call definitions.
  • Many teams rely on ad-hoc methods like Excel sheets to compare prompts, lacking systematic tracking and collaborative evaluation.
  • Identifying bottlenecks in production agents is difficult, hindering targeted improvements.

2. Components of Agent Evaluation

  • The presentation will cover evaluating agents at the tool call level, trajectory level, and across multi-turn conversations.
  • It will also discuss an approach to self-improving agents through iterative evaluation.

3. Tool Calling Evaluation: Ensuring Correct Tool Selection and Argument Passing

  • Agents must choose the right tool call based on context and pass the correct arguments.
  • Evaluation should focus on both aspects: did the agent call the right tool, and did it pass the right arguments?
  • Example: Arise uses its own co-pilot (an insights tool) to troubleshoot application issues and evaluates its traces to identify areas for improvement.
  • Arise's co-pilot architecture follows an orchestration worker pattern, with a planner deciding which tools to call.
  • A high-level view of different agent paths helps pinpoint performance bottlenecks.
  • Example: Analysis of the Arise co-pilot revealed poor performance on search-related questions, specifically in passing the correct arguments to the search tool, even when the correct tool was called.

4. Trajectory Evaluation: Assessing the Order of Tool Calls

  • Trajectory evaluation focuses on whether the agent calls tools in the correct order across a series of steps.
  • Incorrect tool call order can lead to increased token usage and incorrect outputs.
  • Teams should analyze entire traces to evaluate the correctness of the tool calling order.
  • Key Question: Is the agent consistently executing the same set of steps to complete a task, or does it deviate?

5. Multi-Turn Conversation Evaluation: Maintaining Consistency and Context

  • Agent interactions are often multi-turn, requiring the agent to maintain context from previous turns.
  • Evaluation should consider factors like consistency in tone, avoidance of repetitive questions, and effective context tracking.
  • Key Questions: Is the agent consistent in tone? Is it asking the same questions repeatedly? Does it keep track of context from previous turns?
  • Example: Evaluating a multi-turn conversation to ensure the agent correctly answers all questions and maintains context throughout the interaction.

6. Self-Improving Agents: Iterating on Both Agent and Eval Prompts

  • It's crucial to evaluate the agent correctly to identify failure cases and refine prompts.
  • However, the evaluation prompts themselves should also be iteratively improved.
  • Consistently check eval outcomes to identify instances where the eval incorrectly labeled an output as wrong.
  • Iterate on eval prompts, build a golden data set, and refine it continuously.
  • There are two iterative loops: one for agent application prompts and one for eval prompts.
  • Both loops are essential for creating a good product experience.

7. Conclusion: Key Takeaways and Resources

  • Agent evaluation is a multi-faceted process that requires careful consideration of tool calling, trajectory, and multi-turn interactions.
  • Iterating on both agent and eval prompts is crucial for building high-quality agents.
  • Resources: Arise Phoenix (open-source) and Arise Exists.

8. Notable Quotes

  • "Building agents is incredibly hard."
  • "...there really is kind of another loop going on here which is about improving the eval prompts..."

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.