How memory makes AI agents more effective

Google Cloud TechAbout 4 min readSep 14, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Sessions: Short-term memory or the current state of interaction with an agent.
  • Working Memory: Memory used during a specific task.
  • Long-Term Memory: Persistent storage of user attributes, order history, and preferences.
  • Context Engineering: Designing and managing the information provided to an agent to improve its performance.
  • Model Context Protocol (MCP): A protocol for managing memory and state within AI systems.
  • Semantic Search: Searching for information based on meaning and similarity rather than exact keywords.

Deep Dive into Sessions and Memory for Agents

This episode of Real Terms for AI focuses on using sessions, short-term memory, and long-term memory to enhance the performance of AI agents, particularly in application development and data management. The discussion uses a pet shop example to illustrate how these concepts can be applied.

The Pet Shop Agent Example

The scenario involves a user interacting with a pet shop agent to find products for their pets. The initial flow involves the agent querying a product catalog database based on the user's immediate question (e.g., "What types of toys are great for kittens?"). The LLM uses semantic search to find relevant results.

The Importance of Memory

The key problem highlighted is that the LLM is only aware of the user's current question and intent, not their past interactions, preferences, or other relevant information. Memory is crucial to address this limitation.

Types of Memory

The discussion revisits three types of memory:

  1. Working Memory: Memory used during a specific task.
  2. Short-Term Memory (Session): Captures the current state of the interaction.
  3. Long-Term Memory: Stores persistent information about the user, such as order history and preferences.

Leveraging Long-Term Memory

Long-term memory, typically stored in a database, contains information like:

  • User and pet details
  • Order history (recent and past)
  • User-defined preferences

When a session starts, a subset of this information is retrieved from long-term memory and loaded into short-term memory (the session). This context is then provided to the LLM along with the user's question, enabling more relevant and personalized responses.

Example: If the user has a preference for a specific brand of cat food stored in long-term memory, the agent can use this information to filter search results and provide recommendations aligned with that preference.

Updating Long-Term Memory with Short-Term Information

The episode addresses how to incorporate new information learned during a session into long-term memory.

Scenario: The user informs the agent that they recently got a kitten, a fact not yet stored in the database.

Process:

  1. At the end of the session, the new information is sent to a processor.
  2. The processor (potentially using an LLM) summarizes the key updates (e.g., "User has a new kitten").
  3. This summarized information is then used to update the user's profile in the long-term memory database.

This ensures that the agent is aware of the most current information about the user in future interactions.

User Control and Compliance

The importance of providing users with control over their data and preferences is emphasized.

Mechanism: An interface is provided where users can manage their preferences (e.g., "I don't want this type of food anymore").

Process: The same flow is used: information from short-term memory is sent to the processor, which updates the long-term memory database to reflect the user's changes.

Memory for System Improvement

Memory can also be used to improve the agent's performance and tool selection.

Scenario: The agent uses the wrong tool for a specific task.

Solution: The system can store information about which tools should not be used in certain situations. This prevents the agent from making the same mistake in the future.

Example: Storing that a specific tool shouldn't be used for "orders" even if it sounds like it should be.

Model Context Protocol (MCP)

MCP is introduced as a tool for managing memory and state within AI systems.

Functionality:

  • Creating and updating memories
  • Asking the agent questions about how to improve performance
  • Controlling which memories are sent to specific tools or sub-agents

MCP facilitates the handoff of information between different components of the system, similar to how data is transferred between databases, APIs, and services in traditional software engineering.

Summary and Key Takeaways

The episode concludes with a summary of the key concepts:

  • Memory and state are crucial for improving agent performance.
  • Short-term memory (sessions) captures the current interaction.
  • Long-term memory stores persistent user information.
  • Memory can be used to improve tool selection and prevent errors.
  • MCP provides a framework for managing memory and state within AI systems.

The presenters emphasize that context engineering is essentially software engineering, focused on providing the right information to the agent to achieve high-quality results.

Conclusion

By effectively utilizing short-term and long-term memory, AI agents can become more personalized, efficient, and accurate. The integration of user control and system improvement mechanisms further enhances the overall user experience and ensures responsible data management. The use of tools like MCP can streamline the management of memory and state, making it easier to build and maintain complex AI systems.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.