The Agent Factory - Episode 2: Multi-Agent Systems, Concepts & Patterns

Google Cloud TechAbout 5 min readJul 15, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Software 3.0, Context Engineering (Isolating, Persisting, Compressing), Gemini CLI, A2A (Agent-to-Agent) Protocol, Multi-Agent Systems (Supervisor, Deterministic Flows, Dynamic/Swarm), Router Supervisor Pattern, Parallel Supervisor Pattern, Sequential Flow, Circular Flow, Observability (Logging, Tracing), Cloud Monitoring, OpenTelemetry, Traces, Spans.

Software 3.0 and Context Engineering

The discussion begins with the concept of "Software 3.0," a paradigm shift where programming involves natural language prompts and LLMs act as the operating system. This contrasts with Software 1.0 (traditional encoding) and Software 2.0 (training neural networks). A key aspect of Software 3.0 is "Context Engineering," which acknowledges that the bottleneck is not the model's intelligence but managing the limited context window (analogized to RAM).

Context engineering involves three core strategies:

  • Isolating: Passing only relevant data and schema to sub-agents to reduce noise and distractions.
  • Persisting: Leveraging past conversations and retrieving memories based on relevance and recency.
  • Compressing: Using summarization to shorten the context and extracting important facts.

The focus is shifting from designing perfect prompts to building sophisticated runtime architectures around the model.

Gemini CLI

Google released the Gemini CLI, a tool that brings the power of Gemini directly to the command line. It can understand codebases, fix bugs, help with Git commands, create slide decks, organize photos, and even implement YouTube tutorials. It can connect to local tools and MCP servers, expanding its capabilities. The Gemini CLI is presented as an example of a specialized agent.

A2A Protocol and Multi-Agent Systems

Google donated the A2A (Agent-to-Agent) protocol to the Linux Foundation, aiming to create a standard, open way for AI agents to communicate, regardless of their origin. Red Hat and Google released a new A2A sample featuring Python and Java agents using LChain classic, Longchain4j, and ADK. The A2A inspector, a web-based open-source tool, helps developers debug A2A servers and implement the protocol.

Multi-Agent Systems: When and Why

The discussion transitions to multi-agent systems, highlighting that single-agent systems struggle with complexity and multiple tools/domains, leading to performance drops. Multi-agent systems are more resilient and handle complexity better.

A rule of thumb is to consider a multi-agent system when a task requires multiple distinct skills. This allows for specialization and built-in quality control (e.g., an agent reviewing another's work).

Multi-Agent System Architectures

Several architectural patterns for multi-agent systems are presented:

1. With Supervisor

  • Router Supervisor: The supervisor agent coordinates tasks and delegates them to specialized worker agents one by one. The supervisor interacts with the team one by one. The order of interaction is not predefined and adapts to the situation.
    • Example: A supervisor agent assigns a research task to a research agent, who gathers information. The supervisor then passes the information to a writer agent for summarization. The supervisor can then ask the writer to make changes.
  • Parallel Supervisor: The supervisor delegates tasks that can be run simultaneously, speeding up the process.
    • Example: A market analyst agent uses a supervisor to ask a data agent to pull sales figures, a news agent to summarize recent press, and a social media agent to gauge public sentiment, all concurrently.
    • Challenge: Compartmentalization, where the supervisor may not send critical information to all agents.

2. Deterministic Flows

  • Sequential Flow: Agents operate one after another in a predefined sequence, like an assembly line.
  • Circular Flow: Allows for iterative refinement.
    • Example: A coder agent passes code to a tester agent. If the test fails, the coder agent receives the code back for revision. This cycle repeats until the tests pass.

3. Dynamic/Swarm Pattern

  • All-to-all communication model, potentially chaotic.
  • Difference from Supervisor: Control vs. Collaboration. Supervisor model has a clear boss, while the swarm model emphasizes flexibility. If one agent gets stuck, another can jump in to help.
  • Downside: Potential chaos, requiring careful design to keep agents on track.

Each pattern has trade-offs in terms of control, speed, and complexity, depending on the problem being solved.

Demonstration of Router Supervisor Pattern

A small multi-agent system is demonstrated, showcasing the router supervisor pattern. This system has an orchestrator/router and three specialized agents: a business analyst, a data engineer, and a BI engineer. The goal is to answer business questions based on data in BigQuery.

  • The business analyst analyzes the question and proposes solutions.
  • The data engineer constructs a SQL query to retrieve the data.
  • The BI engineer builds a chart.

The system demonstrates dynamic behavior, as changes to the chart request are directed only to the BI engineer, without involving the other agents.

Observability and Debugging

A question from the developer community addresses the challenge of debugging multi-agent systems deployed in the cloud. The key is observability: logging and tracing every decision point.

This includes logging:

  • Incoming tasks
  • How the LLM breaks down tasks into subtasks
  • Which worker agent is assigned to each subtask and why
  • The supervisor's reasoning for choosing a specific worker

Tools like Cloud Monitoring and libraries like OpenTelemetry can help with logging and observation. Google Cloud Trace Explorer is shown as an example, displaying traces (end-to-end flows) composed of spans (subtasks).

Conclusion

The podcast provides a detailed overview of multi-agent systems, covering their benefits, architectural patterns, and debugging strategies. It emphasizes the importance of context engineering and observability in building effective and resilient agentic systems. The discussion highlights the shift towards more sophisticated runtime architectures and the need for specialized agents to handle complex tasks.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.