Hands-on AI workshop: Graph RAG, Memory & Multimodal Agents

By Google Cloud Tech

Share:

Key Concepts

  • Multimodal Data Ingestion: Processing unstructured data (images, video, text) into structured formats.
  • Graph RAG (Retrieval-Augmented Generation): A search methodology that traverses graph relationships (nodes and edges) to provide context-aware results, surpassing traditional vector-only RAG.
  • ADK (Agent Development Kit): An open-source framework for building, orchestrating, and deploying multi-agent systems.
  • Google Cloud Spanner: A unified database solution used for storing both relational graph data and vector embeddings.
  • Semantic & Hybrid Search: Using vector embeddings for meaning-based retrieval (semantic) and combining them with keyword matching (hybrid) for precision.
  • Agent Memory: Distinguishing between Short-term memory (session state/scratchpad) and Long-term memory (persistent storage via Memory Bank).
  • Sequential Agent Workflow: A pattern where agents (Upload, Extraction, Spanner, Summary) execute tasks in a specific, dependent order.

1. System Architecture & Multi-Agent Pipeline

The lab focuses on building a "Survivor Network" application that matches survivors' needs (e.g., medical, technical) with available skills.

  • Solution Stack:
    • Data Layer: Google Cloud Spanner (Graph + Vector).
    • Intelligence Layer: Gemini Pro/Flash models for reasoning and multimodal extraction.
    • Orchestration Layer: ADK (Agent Development Kit) to manage agent logic and tool usage.
    • Deployment: Cloud Run (serverless) and Agent Engine (managed runtime).

2. Step-by-Step Methodology

  1. Environment Setup: Authenticate via gcloud, clone the repository, and use uv (a Rust-based package manager) to sync dependencies.
  2. Database Configuration: Initialize Cloud Spanner to store graph relationships (survivors, locations, skills, needs).
  3. Embedding Generation: Create text embedding models directly within Spanner using SQL. This avoids the overhead of external notebooks by parallelizing vector generation inside the database.
  4. Tooling & Agent Logic:
    • Service Layer: Implement SQL queries for graph traversal.
    • Tooling Layer: Wrap services into tools that the agent can invoke.
    • Agent Layer: Define the "Root Agent" with system prompts and a toolbox.
  5. Multimodal Processing: Use a Sequential Agent workflow:
    • Upload Agent: Sends data to Google Cloud Storage (GCS).
    • Extraction Agent: Uses Gemini 2.5 to parse multimodal inputs (video/images) into structured entities.
    • Spanner Agent: Updates the graph database with new findings.
    • Summary Agent: Generates a concise summary of the process.

3. Key Arguments & Technical Insights

  • Graph RAG vs. Traditional RAG: Traditional RAG retrieves documents based on semantic similarity. Graph RAG traverses relationships (e.g., "Who has the skill to treat this specific injury?"), providing deeper context and more accurate, actionable results.
  • Hybrid Search: Combining semantic search (meaning) with keyword search (exact codes/locations) is recommended for production systems to balance accuracy and specificity.
  • Efficiency of In-DB Models: By defining models (Gemini/Embeddings) inside Spanner, developers reduce latency and complexity compared to managing external API calls from Python code.
  • Callbacks: The use of after_agent callbacks allows for custom logic (like logging or memory extraction) at specific points in the agent lifecycle.

4. Notable Quotes

  • "Model as a brain, tools in the toolbox: the model does the reasoning and picks the tool to solve the problem." — Annie
  • "Cosine distance is optimal for RAG because it focuses on semantic alignment rather than magnitude distance." — Deborah Io
  • "Short-term memory is the existing state (session dictionary), while long-term memory (Memory Bank) allows insights to be shared across multiple sessions." — Deborah Io

5. Data & Research Findings

  • Embedding Dimensions: The text embedding model used in the lab produces a 768-dimensional vector.
  • Memory Bank: Uses Gemini to extract facts from conversations, which are then stored as persistent text summaries, allowing the agent to "remember" past interactions across different sessions.

6. Synthesis/Conclusion

The session demonstrates that building sophisticated, agentic applications requires a robust orchestration framework (ADK) and a unified data strategy (Spanner). By moving from simple RAG to Graph RAG and implementing multimodal sequential pipelines, developers can create systems that not only search data but understand complex, real-world relationships. The transition from in-memory session management to persistent Memory Bank services is the critical step for moving from a prototype to a production-ready, context-aware AI application.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video