How to build AI agents with memory
By Google for Developers
Key Concepts
- ADK (Agent Development Kit): Google's framework for building AI agents.
- Vertex AI Memory Bank: A managed service for persistent, long-term memory for AI agents, part of Agent Engine.
- Volatile Memory (Short-term/In-memory): Memory that is quick and easy to use but is lost once a conversation session ends. Primarily for development.
- Persistent Memory (Long-term): Memory that stores information across sessions, forming an agent's long-term knowledge base. Essential for production.
- Session: A container holding all information for the current conversation, including prompts, agent interactions, and state.
- Memory Service: ADK's abstraction layer for orchestrating calls to underlying memory storage systems.
- Base Memory Service: A common interface in ADK that defines methods for generating and retrieving memories.
add_session_to_memory: A method used by memory services to extract and persist information from a session.search_memory: A method used by memory services to retrieve relevant context based on a user ID and query.- Memory Extraction: The process of identifying and pulling out meaningful information from a conversation.
- Consolidation: The process of combining new memories with existing ones, deduplicating, updating, and curating information over time.
- Scope: An isolation key used in Vertex AI Memory Bank (typically
app_name+user_id) to ensure memories are retrieved and consolidated correctly for specific contexts. - Agent Engine: An umbrella platform within Vertex AI that includes Memory Bank and other agent-related products.
- Express Mode: A Vertex AI mode that allows users to get started with an API key without needing a full GCP project.
- Tokens: Units of text used by LLMs, directly impacting cost and processing.
- Similarity Search: A retrieval method that finds memories semantically similar to a given query.
- Direct Memory Source: A mechanism to upload pre-extracted facts to Memory Bank for consolidation, allowing modular access to Memory Bank's features.
- Callbacks: Functions executed automatically at specific points in an agent's lifecycle (e.g., before or after an agent interaction).
- Preload Memory Tool: An ADK-built-in tool that always runs before an agent's execution, adding relevant memories to the system instructions.
- Load Memory Tool: An ADK-built-in tool that the agent decides to invoke if it determines memory retrieval is necessary.
- Managed Topics: Predefined categories for memory extraction provided by Google (e.g., personal information, user preferences, key conversation events, task outcomes, explicit instructions).
- Custom Topics: User-defined categories for memory extraction, allowing developers to specify labels, descriptions, and few-shot examples for what constitutes "meaningful" information.
- TTL (Time-To-Live): An expiration period for memories, after which they are no longer available.
- Agent Engine SDK: A direct SDK for interacting with Vertex AI Memory Bank, offering more granular control and transparency than ADK's abstraction.
Introduction: The Problem and Solution
The video highlights that a significant and costly issue in AI agents is their lack of memory, leading to frustrating user experiences and increased operational costs as agents repeatedly learn the same information. The solution presented is to build agents that can remember using Google's Agent Development Kit (ADK) and Vertex AI Memory Bank. This approach aims to make agents smarter and more cost-effective over time by enabling them to store and retrieve information across conversations.
Core Memory Concepts in ADK
Memory in AI agents is defined as the ability to store and retrieve information from conversations across sessions. This is conceptualized as:
- Short-term Memory (Volatile Memory): Pertains to the current chat session. It's quick and easy but disappears once the conversation ends. ADK provides the in-memory memory service for this, which is suitable for development and testing but not recommended for production.
- Long-term Memory (Persistent Memory): Acts as the agent's enduring knowledge base, retaining information across multiple sessions. For this, ADK offers options like the managed Vertex AI Memory Bank, integration with other third-party solutions, or custom database implementations.
Memory Workflow
A four-step workflow is outlined for storing and retrieving memory in ADK agents:
- Initialize your memory service: Set up the chosen memory service (e.g., in-memory or Vertex AI Memory Bank).
- Add your session data to your memory service: The current conversation's data (prompts, agent communications, state) is sent to the memory service, which then decides what information to persist and consolidates this raw data.
- Make your memory service accessible to your agent: Configure the agent to interact with the memory service.
- Use ADK built-in tools or APIs to retrieve context: Employ tools like the
preload_memory_toolto automatically fetch memories or thesearch_memoryAPI to retrieve context as needed.
ADK Memory Services: In-Memory vs. Vertex AI Memory Bank
Kimberly Milin, Tech Lead for Vertex AI Memory Bank, details the two primary memory services within ADK:
1. ADK In-Memory Memory Service
- Definition: Requires no setup; data is stored locally in a dictionary on the computer/VM, keyed by app name and user ID.
- Storage: Saves all raw, turn-by-turn conversation events without extraction or LLM processing. This results in very verbose, high-token usage storage.
- Retrieval: Uses simple keyword overlap for
search_memory. If a query word (e.g., "bicycle") matches words in stored memories, those verbose memories are returned. - Limitations:
- Data is lost upon restart or VM change.
- No information extraction or consolidation, leading to high token usage and duplication if the same information is added multiple times.
- Memories are often not actionable due to verbosity and lack of filtering.
- Use Case: Primarily for proof-of-concept and early development stages where quick testing of memory functionality is needed, not for production.
2. Vertex AI Memory Bank Service
- Definition: A managed layer on top of Agent Engine Memory Bank, designed for persistent, high-quality memory.
- Setup: Requires creating an Agent Engine (an umbrella for Vertex Agent products). Can use a GCP project or Vertex AI Express Mode (API key only). Configuration allows specifying the LLM (e.g., Gemini 2.5 Flash for instruction following).
- Storage (
add_session_to_memory):- Makes asynchronous, non-blocking HTTP calls to Memory Bank. Memory generation happens in the background, meaning the client doesn't wait for completion.
- Sends turn-by-turn conversation, but Memory Bank extracts meaningful information and does not persist everything.
- Uses
scope(e.g.,app_name+user_id) as an isolation key for memories, ensuring consolidation and retrieval are context-specific.
- Retrieval (
search_memory):- Uses the
app_nameanduser_idto build the scope key. - Retrieves memories using similarity search based on the user's query.
- Packages the response into an ADK
SearchMemoryResponseobject. - Note: The ADK response might not contain all information available directly from the Agent Engine SDK, which offers more transparency.
- Uses the
- Advantages:
- Intelligent memory extraction (only meaningful information is saved).
- Consolidation (deduplication, updating, and curating memories over time). This makes it a "heuristic memory" that evolves.
- Significantly reduces token usage compared to raw storage.
- Asynchronous generation prevents client latency.
- Designed for production use cases.
Memory Generation: Extraction and Consolidation
Vertex AI Memory Bank's memory generation involves two key automated methods:
- Memory Extraction: Information is extracted from the conversation only if it matches Memory Bank's definition of "meaningful." This definition can be based on managed topics (Google-defined) or custom topics (user-defined).
- Consolidation: New extracted information is compared with existing memories in the corpus. Duplicative, contradictory, or complementary information is combined, ensuring no redundant data and that memories evolve with new context. This process self-curates memories.
Example: Uploading "I like the idea of getting my niece a bike for her third birthday" and then later "I got my niece a red bike for her third birthday" would result in the existing memory being updated rather than a new, duplicative memory being created.
Asynchronous Nature: Memory generation in Memory Bank is non-blocking and happens in the background. This is because it involves multiple LLM calls, which can be latency-intensive. Memories are typically needed for the next turn, not the current one. For blocking generation, the Agent Engine SDK can be used directly.
Direct Memory Source: Allows agents to provide pre-extracted facts directly to Memory Bank for consolidation, enabling modular use of Memory Bank's features.
Automating Generation:
- While the ADK runner doesn't automatically save sessions to memory, developers can call
save_session_to_memoryorgenerate_memoriesdirectly. - A more automated approach is to use callbacks (e.g.,
after_agent_call) to trigger memory generation at the end of each agent interaction. This requires refreshing the session object if callingadd_session_to_memorydirectly, but callbacks provide access to an already populated session. - To avoid sending duplicative information, one can send only the latest turn's content from the callback context using the Agent Engine SDK, though this might lose context from prior turns.
Memory Retrieval Strategies
ADK offers several ways to retrieve memories and inject them into an agent's context:
-
ADK Preload Memory Tool:
- Functions more like a callback than a tool, as it's not invoked by the model.
- It always runs before the agent is executed.
- Dynamically includes generated memories directly into the agent's system instructions.
- Example: An agent, after being told "I plan to get her a doll for Christmas," can later retrieve this information from memory when asked, "Can you remind me what I got my niece for Christmas?" even in a new session.
-
ADK Load Memory Tool:
- Similar to the preload tool but with a key difference: the model needs to decide that loading memory is necessary.
- The system instructions inform the agent that memory is available and can be looked up. The agent then invokes the tool if relevant.
- Example: If asked "What should I get my niece for Christmas?", the agent might invoke the tool. If asked a simple "Hi," it likely won't, as memory isn't useful.
-
Custom Callbacks:
- Provides more granular control over memory retrieval.
- Allows developers to define their own scope keys, customize how information is formatted in the prompt, and directly call Memory Bank.
- Example: A custom callback using only
user_idas a scope key can retrieve memories specific to that user, even if ADK's default scope (user ID + app name) was different. This requires generating memories with the same custom scope.
Customizing Memory Bank Behavior
Memory Bank offers extensive customization to align with specific business needs:
1. Customizing Memory Extraction Topics
- Managed Topics: By default, Memory Bank persists information categorized as personal information, user preferences, key conversation events, task outcomes, and explicit instructions to remember/forget.
- Customizing Managed Topics: Users can specify a subset of managed topics to persist (e.g., only user preferences, ignoring key conversation events). This configuration can be scope-specific.
- Custom Topics: For use cases not covered by managed topics, developers can define their own custom topics.
- This involves providing a custom label, description, and few-shot examples (sample conversations with expected memory outcomes, including "no op" examples for non-meaningful information).
- Example: Creating a custom topic to extract and condense "feedback for your business" (e.g., "I think that the coffee shop should offer more milk options such as almond milk" from customer feedback). This allows Memory Bank to act as a data mining tool for specific information.
Time-To-Live (TTL) for Memories
- Default TTL: A default TTL can be set for all generated memories when setting up Memory Bank (e.g., 30 days). This is crucial because memories are agentically managed and mutated by Memory Bank, not explicitly created by the user.
- Granular TTL: For more control, TTL can be defined per operation that created the memory. This allows for scenarios where, for instance, the TTL is set only when a memory is created but not refreshed when it's updated through consolidation.
- Example: A memory created with a one-year TTL will retain that expiration date even if it's updated multiple times through consolidation, as long as the granular TTL configuration specifies not to refresh on update.
Q&A: Deep Dive and Best Practices
-
Why use Memory Bank instead of just a database?
- Databases store raw, verbose dialogue (like the in-memory service), leading to high token usage and extraneous information.
- Memory Bank saves only meaningful information and uses consolidation to deduplicate and curate data, storing only what's necessary for the future.
- It performs processing at storage time to reduce information, shifting the burden from retrieval time.
-
Does instructing an agent to "remember" something make Memory Bank remember it?
- The agent orchestrates calls to tools/services. If the agent has a tool to extract and send memories to Memory Bank, then the instruction would be helpful.
- However, Memory Bank is ultimately responsible for extraction. It's more effective to define such instructions as a custom topic within Memory Bank's configuration so Memory Bank itself knows what type of information to persist.
-
Can Memory Bank be used with models other than Gemini (e.g., Olama, Mistral)?
- No, Memory Bank currently only supports Gemini models. The specific Gemini model (e.g., Gemini 2.5 Flash) is provided during Memory Bank setup.
Conclusion and Main Takeaways
The video emphasizes the critical role of memory in building personalized, intelligent, and cost-effective AI agents. It provides a comprehensive guide to leveraging Google's ADK and Vertex AI Memory Bank to achieve this. Key takeaways include:
- Memory is essential: It prevents agents from repeatedly learning the same facts, reducing frustration and cost.
- Persistent memory is key for production: Vertex AI Memory Bank offers a robust, managed solution for long-term memory.
- Intelligent processing: Memory Bank excels through automated memory extraction and consolidation, significantly reducing verbosity and token usage compared to raw storage.
- Flexible retrieval: ADK provides built-in tools (
preload_memory_tool,load_memory_tool) and allows for custom callbacks for tailored memory access. - Extensive customization: Developers can define what constitutes "meaningful" information through managed or custom topics and manage memory lifecycles with TTL settings.
- Managed services for complexity: Memory Bank handles the underlying complexities of memory generation and retrieval, making it ideal for high-quality, production-grade agents.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development