Build multi-agent AI A2A + Cloud Run | Hands On AI (Part 2)

By Google Cloud Tech

Share:

Key Concepts

  • A2A (Agent-to-Agent Protocol): A universal protocol enabling communication between remote agents, allowing them to discover each other and share capabilities via "Agent Cards."
  • Agent Card: A discoverable JSON file at a well-known path (/.well-known/agent-card.json) that acts as an agent's resume, detailing its skills, capabilities, and security/authentication requirements.
  • ADK (Agent Development Kit): The framework used to build, orchestrate, and manage the lifecycle of AI agents.
  • Callbacks: Hooks in the agent lifecycle (e.g., before_agent, after_tool) used to inject custom logic, such as security filtering or state management.
  • Plugins: Global versions of callbacks that apply logic across all agents within an ADK runtime.
  • Agent State: A dictionary-based storage mechanism used to track specific, concrete data points (e.g., "last_summoned_familiar") across agent interactions.
  • MCP (Model Context Protocol): A standard for connecting agents to external tools and data sources.

1. Agent-to-Agent (A2A) Protocol

A2A facilitates communication between agents deployed on different endpoints (e.g., Cloud Run, GKE, or local environments).

  • Discovery: The "Summoner" (orchestrator) agent retrieves the Agent Card of sub-agents to understand their skills and how to authenticate with them.
  • Security: A2A supports security schemes like API keys or OAuth, ensuring only authorized agents can invoke specific services.
  • Implementation: Developers enable A2A by importing to_A2A from the ADK library and passing the agent instance, port, and public URL.

2. Callbacks and Plugins

Callbacks allow developers to intercept the agent lifecycle to implement custom logic.

  • Lifecycle Hooks: before_agent_callback, after_agent_callback, before_tool_callback, and after_tool_callback.
  • Real-world Application:
    • Security: Using "Model Armor" in a before_agent callback to filter malicious inputs or prevent data leakage.
    • Efficiency/Throttling: Implementing a "cooldown" logic in a before_agent callback to prevent service spamming.
  • Plugins vs. Callbacks: Callbacks are defined at the individual agent level, whereas Plugins are defined at the ADK runtime level, applying rules globally to all agents managed by that runner.

3. Agent State and Memory

  • Memory: Refers to the broader conversational history or semantic context of an interaction.
  • State: Refers to specific, structured data (key-value pairs) extracted from the conversation.
  • Methodology: By using an after_tool_callback, the Summoner agent updates a state variable (last_summon) whenever a familiar is delegated a task. This allows the agent to maintain context without re-parsing the entire conversation history.

4. Orchestration Framework

The lab demonstrates a hierarchical routing pattern:

  1. Summoner Agent (Orchestrator): Uses Gemini 2.5 Flash to analyze incoming prompts and semantically select the best sub-agent (Fire, Water, or Earth) for the task.
  2. Sub-Agents: Specialized agents (Sequential, Parallel, or Loop) that execute specific tasks.
  3. Communication: The Summoner delegates tasks to these remote agents via A2A, treating them as if they were in the same local runtime.

5. Notable Quotes

  • "Agent cards are essentially the resumes for agents... think about it like the Instagram profile for an agent." — Io
  • "A2A is better for agent-to-agent communication, while MCP is for tool discovery and tool communication. They are collaborative, not competitive." — Io

6. Synthesis and Conclusion

The session successfully demonstrated the transition from local agent development to a distributed, multi-agent architecture. By leveraging A2A, developers can decouple agents into independent microservices, enhancing scalability and maintainability. The use of Callbacks and Plugins provides a robust mechanism for enforcing organizational policies (like security and rate limiting) across complex systems. Finally, the integration of State management ensures that agents remain context-aware, enabling sophisticated, multi-step workflows like the "boss fight" scenario, where the orchestrator must strategically select and manage familiars based on the specific weaknesses of the opponent.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video