The Agent Factory - Episode 9: Agent evaluation with ADK & Vertex AI
By Google Cloud Tech
Key Concepts
Agent evaluation, LM evaluation, traditional testing, deterministic vs. variable systems, full-stack evaluation, task success rate, output quality, chain of thought, tool utilization, memory and context retention, offline evaluation, online evaluation, ground truth checks, LLM as a judge, human in the loop, calibration loop, ADK web, Vertex AI, synthetic data generation, three-tier testing strategy (unit, integration, end-to-end), multi-agent systems, agent hand-off, collaboration, communication efficiency, cost-scalability trade-off, benchmark integrity, subjective attribute evaluation.
What is Agent Evaluation?
Agent evaluation is a comprehensive assessment of an agent's performance, going beyond simply checking if the final answer is correct. Unlike traditional software testing, which is deterministic (same input, same output), agent behavior is variable, leading to different outcomes even with the same prompt. It's not just about the final output, but also about the system-level behavior, including:
- Autonomy: The agent's ability to work independently.
- Multi-step reasoning: The agent's ability to break down tasks into logical steps.
- Tool usage: How effectively the agent uses available tools.
- Handling unpredictable situations: How the agent responds to unexpected events.
Difference from LM Evaluation: LM evaluation (like MMLU) tests knowledge with static Q&A, similar to a school exam. Agent evaluation is more like a job performance review, focusing on the agent's ability to use tools, recover from errors, and maintain consistency across multiple turns. A great model doesn't guarantee a great agent if the agent can't call APIs properly.
Key takeaway: Traditional testing is for deterministic logic, LM evaluation is for general model capabilities, and agent evaluation is for system-level task effectiveness.
Full-Stack Approach to Agent Evaluation
A full-stack approach to agent evaluation involves measuring everything, including:
- Final Outcome:
- Task Success Rate: Did the agent achieve its goal?
- Output Quality: Was the output coherent, accurate, and safe? Did it avoid hallucination and bias?
- Agent's Chain of Thought:
- Did the agent break the task into logical steps?
- Was the reasoning consistent?
- Tool Utilization:
- Did the agent pick the right tool?
- Did it pass the correct parameters?
- Was the tool usage efficient (avoiding redundant API calls)?
- Memory and Context Retention:
- Can the agent recall the right information when needed?
- Can the agent resolve conflicts between new and existing information?
How to Evaluate an Agent in Practice
Two main concepts:
- Offline Evaluation: Pre-production testing against static golden datasets to catch regressions.
- Online Evaluation: Post-deployment monitoring of live user data for drift or running A/B tests.
Popular Methods:
- Ground Truth Checks: Fast, cheap, and reliable for checking basic aspects like valid JSON generation or schema matching. Limitation: Doesn't capture nuances like coherence or factuality.
- LLM as a Judge: Using a strong model to score subjective qualities like plan coherence. Scales well, but evaluation depends on how the LLM was trained.
- Human in the Loop: Domain experts review agent outputs. Most specialized but also slower and more expensive.
Calibration Loop: The best strategy is to combine these methods. Start with human experts to create a small, high-quality golden dataset, then fine-tune an LLM as a judge until its scores align with human expectations. This combines the accuracy of humans with the scalability of LLMs.
Agent Evaluation with ADK and Vertex AI
ADK Web: Agent Development Kit (ADK) web UI is designed for fast, interactive offline development. It facilitates a five-step loop:
- Test the Agent and Define the Golden Path: Input a prompt and observe the agent's response.
- Create a Case: In the eval tab, correct the expected response to match the desired output, creating a golden data set.
- Evaluate the Agent: Select the test case and run the evaluation.
- Find the Root Cause: Use the trace tab to examine the agent's step-by-step reasoning process.
- Fix and Validate the Agent: Modify the agent's code based on the root cause analysis and rerun the evaluation to ensure the test passes.
Example: A product research agent with two tools: getProductDetailsForCustomerFacingInfo and lookupProductInformationForInternalSKUs. The initial instruction is ambiguous, leading the agent to use the wrong tool. By clarifying the instruction ("For customer-facing description, use getProductDetails; for internal data like SKU, use lookupProductInformation"), the issue is resolved.
Vertex AI: For testing at scale and richer metrics (using LLM as a judge), a production-grade platform like Vertex AI is needed. Vertex AI's GenAI evaluation services handle complex qualitative evaluations at scale and produce evaluation outcomes that can be used to build dashboards.
Key takeaway: ADK is for fast inner loop development, while Vertex AI is for production scale-out.
Synthetic Data Generation
Addresses the cold start problem (lack of available data) by using an LLM to create the dataset. A four-step recipe:
- Generate Realistic User Tasks: Ask the LLM to create realistic user tasks.
- Produce Perfect Solutions: Have the LLM act as an expert agent and generate perfect, step-by-step solutions.
- Generate Imperfect Attempts (Optional): Use a weaker model to attempt the same tasks, creating imperfect attempts.
- Score Attempts: Use an LLM as a judge to compare the imperfect attempts against the perfect solutions and score them automatically.
Three-Tier Testing Strategy
- Unit Tests: Testing small pieces of the agent in isolation (e.g., testing a single tool).
- Integration Tests: Testing the entire multi-step journey to ensure all components work together correctly.
- End-to-End Human Review and Multi-Agent Testing: A final sanity check involving multiple agents and feeding results back into the human-in-the-loop calibration loop.
Multi-Agent System Evaluation
In multi-agent systems, evaluating individual agent initiations doesn't provide a complete picture of overall system performance. The focus shifts to whether the whole system gets the job done.
Example: Agent A (customer support) hands off a customer request to Agent B (refund and replacement). If Agent A's task completion score is zero because it doesn't process the refund itself, this is misleading. What matters is whether Agent A can hand off smoothly, share context, and maintain reasonable latency and cost across the entire journey.
Key takeaway: End-to-end evaluation is crucial in multi-agent systems. Evaluation becomes part of the design, and agents may need to emit structured data specifically for evaluation purposes. Multi-agent evaluation can be viewed as a network analysis, focusing on interactions rather than just outcomes.
Open Questions and Challenges
- Cost-Scalability Trade-off: Human evaluation is accurate but slow and expensive. LLM as a judge is faster and scalable but requires tuning to align with human expectations.
- Benchmark Integrity: Test questions leaking into model training data can invalidate scores.
- Subjective Attribute Evaluation: Measuring subjective qualities like creativity, productivity, and humor in agents remains a challenge.
Conclusion
Agent evaluation is a complex but crucial aspect of putting agents into production. It requires a full-stack approach, combining various evaluation methods and tools. As agents become more sophisticated and multi-agent systems become more prevalent, new evaluation frameworks are needed to address the emerging challenges.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development