How to build reliable AI Agents?

Google for DevelopersAbout 6 min readNov 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Agent Observability: The practice of instrumenting, monitoring, and analyzing the behavior of AI agents to ensure their reliability and performance in production.
  • Non-deterministic Systems: AI agents, particularly those powered by LLMs, exhibit unpredictable behavior due to their complex nature, unlike traditional deterministic software.
  • Agent Development Kit (ADK): A framework from Google for building AI agents, providing plugins and callbacks for inspecting agent actions.
  • Phoenix: An open-source observability tool from Arise AI that allows developers to trace, evaluate, and iterate on AI agents.
  • Open Telemetry (OTel): An open-source observability framework used for instrumenting and collecting telemetry data (traces, metrics, logs) from applications.
  • Open Inference: An OTel-compliant specification for GenAI applications, enabling interoperability between tools.
  • Error Analysis: The process of inspecting agent traces to identify logical flaws or unexpected behavior, even when no explicit exceptions are thrown.
  • Evals (Evaluations): Methods for assessing the performance and output quality of AI agents, categorized into offline and online evaluations.
  • Answer Correctness Eval: An evaluation method that uses an LLM as a judge to determine if an agent's output accurately answers a user's question.

Observability for AI Agents: From Prototype to Production

This discussion focuses on the critical challenge of transitioning AI agents from prototype to reliable production applications, emphasizing the role of observability. Ivan, a software engineer at Google working on the Agent Development Kit (ADK), is joined by Aperna, CPO and co-founder of Arise AI, a leader in GenAI observability.

The Critical Need for Observability in GenAI

Aperna highlights that traditional software had deterministic paths, allowing for debugging with logs and step-through debuggers. However, LLM-powered agents are fundamentally different:

  • Non-deterministic Behavior: They exhibit complex, emergent behaviors that can lead to unpredictable failures.
  • Expanded Debugging Scope: Debugging now involves not just code, but also prompts, reasoning processes, and the tools agents utilize.
  • Shift from "Hope" to "Knowledge": Observability provides the necessary instrumentation to understand agent behavior, moving from an uncertain state ("I hope this works") to a confident understanding ("I know how this works"). This is crucial for building trustworthy AI agents.

Ivan confirms that ADK was designed with this in mind, offering plugins and callbacks to inspect all agent actions, including LLM and tool calls. This allows developers to easily understand and control the "what" and the "how" of agent operations, focusing on business logic rather than boilerplate code.

Introducing Phoenix: An Open-Source Observability Tool

Phoenix, Arise AI's open-source tool, is presented as a solution to gain deeper insights into agent operations.

  • Functionality: Phoenix enables developers to trace, evaluate, and iterate on their agents.
  • Use Cases: Teams use it to debug performance issues, evaluate agents both offline and online, and iterate rapidly (in minutes, not hours).
  • Deployment: It can run locally on a developer's machine or in the cloud.

Developer Experience: Integrating Phoenix with ADK

The session demonstrates the seamless integration of Phoenix with an ADK-built agent.

  1. Agent Setup: The example uses a financial advisor agent built with ADK, designed to research stocks. It employs a multi-agent system with sub-agents for stock and risk analysis, utilizing tools for fetching stock data and PE ratios. This example is available on the ADK samples GitHub repo.
  2. Phoenix Integration:
    • Account Creation: Teams need to create a Phoenix account.
    • Instrumentation: A few lines of code are required to instrument the agent. This involves importing Phoenix, Google ADK, and Open Inference (which follows Open Telemetry conventions for GenAI).
    • Tracer Configuration: A standard tracer provider is configured to trace the application.
    • Passing the Tracer: The configured tracer is passed directly into the ADK agent's constructor during initialization.
  3. Open Standards and Open Inference: Ivan emphasizes the importance of open standards like Open Telemetry for creating a common language for tools to communicate. Open Inference is designed to be OTel compliant for GenAI applications, facilitating easy data transfer to Phoenix, especially since ADK also emits OTel traces.

Live Demo: Tracing Agent Execution

A live demo showcases the agent's behavior and how traces appear in Phoenix.

  • Initial Query: The agent is asked for a trading strategy for Apple, Google, and Microsoft. The financial coordinator agent initiates data analyst agents to process the request.
  • Live Tracing: As the agent works, all its actions are traced and streamed directly into Phoenix, allowing for real-time observation of live traces.
  • Second Query and Failure: A subsequent query for Google's stock price results in the agent stating it cannot fulfill the request.

Error Analysis: Inspecting Failed Requests

The inability to fulfill the second request prompts an inspection of the agent's internal workings.

  • Problem Identification: The initial instinct is to use print statements within after model call callbacks to log the model's thinking. Phoenix facilitates this by showing all steps taken by the agent.
  • Side-by-Side Comparison: The traces from the successful and failed queries are compared.
    • Successful Query: Shows clear agent calls, tool calls, and LLM calls.
    • Failed Query: Reveals that the agent called the LLM, but there was no subsequent tool call.
  • Logical Bug: This scenario exemplifies error analysis, where the code runs without exceptions, but the bug lies in the logic. The agent failed to call the necessary tool after the LLM interaction.

Debugging at Scale: The Role of Evals

As applications process more data, manual inspection of every trace becomes impractical. Evals are introduced as a solution for debugging at scale.

  • Purpose of Evals: Evals help understand the quality of application outputs.
  • Types of Evals:
    • Online Evals: Critical for applications in production, they evaluate outputs and identify areas for improvement.
    • Offline Evals: Used during experimentation, akin to unit testing, performed before shipping applications.
  • Answer Correctness Eval Demo: The session demonstrates an online evaluation using the "Answer Correctness Eval." This method employs an LLM as a judge to assess if the agent's output accurately answers the user's question.
  • Interpreting Eval Results: The demo shows rows of traces with labels indicating whether the judge found the answer sufficient. An "incorrect" example reveals an LLM-generated explanation from Phoenix stating the agent couldn't retrieve stock ticker data, possibly due to a missing tool. This provides a clear starting point for debugging.

Synthesis and Conclusion

The collaboration between ADK and Phoenix highlights the power of combining a robust agent-building framework with a comprehensive observability tool.

  • ADK's Contribution: Provides a powerful, scalable framework for building AI agents.
  • Arise Phoenix's Contribution: Offers critical visibility for debugging, iterating, and perfecting these agents.
  • Seamless Integration: Adding a few lines of instrumentation code to an ADK agent provides a production-grade debugging experience.
  • Outcome: This combination enables teams to move from a "cool demo" to a reliable, production-ready AI application that can be confidently shipped to users.

Actionable Insights:

  • Developers can run Phoenix locally or utilize Arise AX or Google Cloud Trace.
  • Resources for further exploration include links to ADK documentation, the open-source Phoenix project on GitHub, and the Financial Advisor agent code.
  • Phoenix is a community-driven project, encouraging contributions through GitHub stars, issues, and feedback.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.