AI agent long-term memory with memory bank
By Google Cloud Tech
Key Concepts
- Session Service: Manages active, live conversations and maintains state.
- Memory Service: Acts as a long-term archive (filing cabinet) for cross-session data.
- Semantic Search: A search method that retrieves information based on meaning rather than exact keyword matches (e.g., searching "two-wheeled vehicle" finds "bicycle").
- Agent Engine: The core infrastructure that extracts facts from media/text, generates embeddings, and manages storage.
- Embeddings: Numerical representations of data that allow the system to understand and compare the semantic meaning of text, images, audio, and video.
- Preload Memory Tool: An automated utility that runs at the start of every chat turn to inject relevant historical facts into the agent's context.
1. Architecture: Session vs. Memory Service
The video distinguishes between two critical components for building personalized agents:
- Session Service: Focused on the "here and now." It handles the immediate state of a conversation, allowing users to resume live chats.
- Memory Service: Focused on the "long term." It serves as a persistent repository that survives restarts and spans multiple days or weeks.
2. Memory Service Options
- In-Memory Service: A lightweight, local solution intended for quick testing. It does not persist data across restarts and relies on basic keyword matching.
- What’s AI Memory Bank Service: A production-ready cloud service. It supports semantic search and persistent storage, allowing the agent to "understand" the context of stored information rather than just indexing strings.
3. Configuration and Ingestion Methodologies
To set up the Memory Bank, the developer must configure an Agent Engine with two specific model types:
- Fact Extraction Model: Uses models like Gemini to parse conversations and media files to identify key information.
- Embedding Model: Converts extracted facts into vector embeddings to enable semantic search.
Ingestion Methods:
- Full Session Archiving: At the end of a chat, the
add_session_to_memoryfunction is called. The engine processes the entire history—including text, images, audio, and video references—to extract and store meaningful facts. - Direct Media Upload: Developers can preload files (images, video, audio) directly into the memory bank with accompanying text context, bypassing the need for a prior chat session.
4. Retrieval Framework: The Preload Memory Tool
The retrieval process is automated to ensure the agent remains context-aware without requiring manual logic:
- Execution: The
preload memory toolruns at the start of every user turn. - Process: It reads the new user message, performs a semantic search against the Memory Bank, and retrieves the most relevant historical facts.
- Injection: These facts are automatically injected into the agent's prompt, providing the necessary context for a personalized response.
5. Real-World Application: Multimodal Recall
The video demonstrates a scenario involving a user sharing diverse media:
- Input: A user shares a photo of a historical building, a video of the coast, and an audio note about a specific town.
- Processing: The Memory Bank extracts facts: "likes historical architecture," "enjoys the coast," and "visited [town name]."
- Recall: In a subsequent, unrelated session, the user asks for a cultural destination recommendation. The agent uses the
preload memory toolto retrieve the previously stored facts and suggests a location that aligns with the user's established preferences for history and seaside environments.
6. Synthesis and Conclusion
The series concludes by defining the three-layer memory stack required for a sophisticated, personalized agent:
- Working Memory: Session and state for live interaction.
- Persistent Memory: User profiles that survive restarts.
- Long-Term Memory (Memory Bank): A multimodal archive that uses semantic search to maintain context over extended periods.
By integrating these layers, developers can build agents that are not only consistent but also deeply context-aware, capable of evolving their understanding of a user over days and weeks.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

The Agentic AI Engineer - Benedikt Sanftl, Mutagent
AI Engineer

Building Great Agent Skills: The Missing Manual
AI Engineer

Agents Building Agents - Alfonso Graziano, Nearform
AI Engineer

Your Agent Is Wasting Tokens and You Don't Know It - Erik Hanchett, AWS
AI Engineer

Agent development and AgentOps with BigQuery, ADK, and MCP
Google Cloud Tech

Google’s new AI agent stack from I/O 2026
Google Cloud Tech

Top Open-Source GitHub Projects : AgentsView, Cypress, Chatwoot, Strands Shell & Raven #268
ManuAGI - AutoGPT Tutorials