Building distributed multi-agent systems

Google Cloud TechAbout 3 min readMay 27, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Agents: Systems that use LLMs to reason, select tools, and perform tasks independently.
  • Multi-Agent System (MAS): A modular architecture where specialized agents (Researcher, Judge, Content Builder) collaborate under an Orchestrator.
  • ADK (Agent Development Kit): A framework for building and running AI agents.
  • A2A (Agent-to-Agent) Protocol: An open communication standard allowing agents to discover and collaborate via HTTP, regardless of the underlying framework.
  • Cloud Run: A serverless platform for deploying containerized microservices.
  • Session State: A shared key-value store managed by the Orchestrator to maintain context across agent turns.
  • Quality Gate (Judge Pattern): A deterministic agent that evaluates outputs against a schema to ensure quality before proceeding.

1. Architecture and Methodology

The system follows a modular, distributed microservices architecture. Instead of a monolithic prompt, the "AI Course Creator" uses a squad of specialists:

  • Researcher: Retrieves information using web search tools.
  • Judge: Acts as a programmatic quality gate, returning structured JSON (Pass/Fail) to prevent hallucinations.
  • Content Builder: Formats the validated research into a structured course.
  • Orchestrator: Manages the global session state, coordinates agent communication, and handles the workflow logic.

Methodologies:

  • Loop Pattern: Used between the Researcher and Judge to iterate until quality standards are met (with a max_iterations limit to prevent infinite loops).
  • Sequential Pattern: Used to pass validated data from the research loop to the Content Builder.
  • Callback Hooks: Deterministic functions that trigger after an agent finishes a task to save outputs to the shared session state.

2. Technical Stack and Tools

  • Google Cloud Run: Hosts each agent as an independent container, allowing for independent scaling and failure isolation.
  • UV: A high-performance Python package manager (written in Rust) used for dependency management.
  • Gemini 3 Flash: The default model used for reasoning and tool usage.
  • Agent Cards: JSON files hosted at well-known URLs that act as "business cards," enabling service discovery via the A2A protocol.

3. Key Arguments and Perspectives

  • Autonomy as a Dial: The presenters argue that agent autonomy is not a binary switch. For quality gates (like the Judge), developers should intentionally restrict autonomy to ensure deterministic, predictable behavior.
  • Production Readiness: Moving from a local script to production requires moving away from monolithic prompts toward modular, testable microservices.
  • Human-in-the-loop vs. Programmatic Gates: The speakers emphasize that manual review of LLM outputs is unscalable; programmatic "Judge" agents are essential for production-grade systems.

4. Notable Quotes

  • "An AI agent is a system that uses a model to reason about and select the appropriate tools to achieve a specific goal." — Priya
  • "Multi-agent autonomy is a dial... not a switch." — Scheer
  • "Small, focused agents are easier to evaluate, debug, and scale independently." — Scheer

5. Practical Insights & Troubleshooting

  • Environment Variables: A common point of failure is failing to source environment variables in new terminal tabs during deployment.
  • IAM Permissions: When deploying to Cloud Run, ensure the service account has the necessary roles (e.g., Storage Object Viewer if accessing artifacts).
  • Debugging: Use the ADK Web UI to test agents in isolation before integrating them into the broader orchestration pipeline.
  • Security: While the demo uses allow-unauthenticated for simplicity, production systems must use IAM-based authentication to prevent malicious abuse.

6. Synthesis/Conclusion

The session demonstrated that building a production-ready AI system requires shifting from "fun local experiments" to a robust, distributed architecture. By utilizing the ADK framework, the A2A protocol, and Cloud Run, developers can create modular systems that are easier to debug, scale, and maintain. The core takeaway is the importance of deterministic quality gates and explicit state management to ensure that AI agents deliver reliable, factually grounded results in real-world applications. Future sessions in this series will focus on evaluation (June 30) and security (July 28).

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.