Key Concepts:
- AI Agents, Memory, Working Memory (Short-Term Memory), Long-Term Memory, Sessions, Vertex AI Memory Bank on Agent Engine, Agent Development Kit (ADK), Retrieval-Augmented Generation (RAG), Memory Extraction, Memory Consolidation, Scopes, Agent Engine Runtime.
1. Agent Industry Pulse (Latest Updates):
- Model Releases: Qwen, Google's Gemini Deep Think in Gemini Ultra.
- Evaluation: LangChain released Align Evals within LangSmith for better automated evaluation matching human preference. Developers can tweak evaluators in a playground and see real-time score comparisons to human labels.
- Specialized Models: Zhipu AI open-sourced GLM-4.5, a large language model designed for agents. It was pre-trained on 15 trillion tokens of general data, followed by 7 trillion tokens focused on code and reasoning, and fine-tuned on instructional and domain-specific datasets. GLM-4.5 ranked third and GLM-4.5-Air ranked sixth on benchmarks focused on reasoning, coding, and agentic tasks. It uses reinforcement learning post-training to enhance agent-like capabilities.
- Agent Frameworks: Agent Development Kit (ADK) and Agent-to-Agent Protocol are rapidly evolving. The latest ADK release introduced Plugins for global-level settings of functions like logging, authentication, and caching. A-to-A version 0.3 includes a more stable Python SDK, gRPC support, and enhanced security.
2. The Importance of Memory in AI Agents:
- The Goldfish Challenge: Typical agents lack memory of previous interactions. Each turn is a new start.
- Working Memory (Short-Term Memory) and Sessions: Sessions allow agents to retain information during a single conversation by keeping information within the LLM context window.
- Limitations of Sessions: Context windows are finite, leading to increased costs, slower performance, and potential noise. Sessions are empty at the start of each new conversation.
- Long-Term Memory: Needed to store, process, and retrieve relevant information across multiple sessions. It enables agents to learn from conversations and extract key facts for long-term use.
3. Vertex AI Memory Bank on Agent Engine:
- Managed Service: A managed service that handles the heavy lifting of long-term memory management for AI agents. It agentically extracts facts, deduplicates them, updates them, and provides a retrieval mechanism.
- Integration with ADK: The PreloadMemoryTool in ADK automatically retrieves relevant memories and injects them into the agent's system instructions.
- Open API: Vertex AI Memory Bank provides an open API for integration with other agent frameworks like LangGraph, CrewAI, and Pydantic AI.
4. Vertex AI Memory Bank vs. Retrieval-Augmented Generation (RAG):
- Kimberly Milam (Tech Lead for Memory Bank) explained the differences.
- Databases vs. Memory Systems: Databases are static, with direct read/write requests. Memory systems extract only the most meaningful information for future conversations.
- RAG vs. Memory: RAG uses external knowledge managed separately from the agent (e.g., company tutorials). Memory uses internal knowledge scoped to a particular user or class of users. RAG is managed externally to the agent, while memory is managed based on the conversation and execution of the agent. RAG is generally read-only from the agent's perspective, while memory is both read and written by the agent.
5. Memory Extraction and Consolidation Process:
- Memory Extraction: LLM extracts meaningful information from conversations based on predefined topics: user preferences, user information, key conversation events, and explicit instructions.
- Memory Consolidation: LLM analyzes existing memories and determines whether to add new memories, update existing ones, or delete outdated ones. This process deduplicates information and resolves contradictions.
- Gemini Used: Gemini is used for memory extraction and consolidation.
- Developer Control: Developers can use the CRUD APIs to directly manipulate the contents of Memory Bank.
6. Managing Memory and Sessions:
- Distinction: Sessions are turn-by-turn conversation history (immutable log). Memories are extracted, condensed, and self-contained information.
- Automated Management: Memory Bank automates memory management through consolidation.
- Manual Control: Developers can use CRUD APIs to directly manipulate memories.
7. Evaluating Memory:
- Open Topic: Evaluation is still an open topic, especially considering the differences between academic datasets and production systems.
- Golden Datasets: Developers can define their own golden datasets of cross-session conversations to test Memory Bank's ability to maintain continuity.
8. Efficiency in Production:
- Background Processing: Memory generation should happen in the background to avoid latency issues.
- Retrieval Methods:
- Retrieving all memories for a particular scope (defined by a dictionary of metadata).
- Using similarity search to retrieve the most relevant memories (incurring latency from the embedding API).
- Memory Quality: Focus on extracting high-quality memories to reduce stress on retrieval.
9. Flexibility of Memory Bank:
- Part of Vertex AI Agent Engine: Memory Bank is part of Vertex AI Agent Engine, but it can be used independently.
- Integration: Can be integrated into various environments (GKE, Cloud Run, Colab) and agent frameworks (LangGraph, CrewAI).
10. Community Questions:
- Official Online Community for ADK: GitHub discussion for feature requests and the Agent Development Kit's subreddit for community support.
- Deterministic Agent Routing: Use the
before_model_callbackto intercept agent actions and implement routing logic. - Shared Memory Across a Team: Use scope-based retrieval with a team ID or department ID to create a shared knowledge base.
11. Conclusion:
Vertex AI Memory Bank on Agent Engine offers a powerful solution for adding long-term memory to AI agents. It automates complex tasks like memory extraction and consolidation, providing developers with a managed service that can be easily integrated into various agent frameworks and environments. While evaluation remains an ongoing area of research, the flexibility and control offered by Memory Bank make it a valuable tool for building smarter and more personalized AI agents.
AI summaries can miss context or contain errors. Check important details against the original video.





