Key Concepts
Context engineering, AI agents, prompt engineering, context window, short-term memory, long-term memory, RAG (Retrieval Augmented Generation), tool calling, context isolation, multi-agent systems, context summarization, context routing, context staging, context formatting, context trimming, context poisoning, context distraction.
Short-Term Memory
- Definition: Retains recent interactions within a defined context window.
- Implementation: Can be implemented using simple memory (within NAN) or external databases like Postgres.
- Example: An AI agent remembers a number provided in a previous interaction.
- Technical Detail: The context window is measured in tokens. Token usage can be monitored within the model settings (e.g., OpenAI chat model).
- Limitation: Only remembers the last few messages defined within the context window.
Long-Term Memory
- Definition: Stores information for future use.
- Implementation: Can be implemented using Google Docs, Google Sheets, Airtable, databases, or RAG systems.
- Process:
- AI agent receives a chat message.
- AI agent decides if the message contains information useful for the future.
- If so, the information is saved to long-term memory.
- Example: An AI agent saves the user's name and platform they use to a Google Doc.
- Customization: AI agent can categorize the type of memories.
- Integration: Long-term memory can be integrated into a RAG system.
- System Message: The system message defines the overall behavior and operating procedure that the agent should follow, including how to use the long-term memory.
Context Expansion via Tool Calling
- Definition: Dynamically injecting information into the context by using external tools.
- Example: An AI agent uses a Perplexity tool to get the biggest AI news items of the week.
- Process:
- AI agent receives a message that requires external information.
- AI agent chooses to call a tool (e.g., Perplexity).
- The tool's response is dynamically injected into the context.
- Limitation: Context can get polluted quickly if not managed carefully.
- Token Usage: Multiple tool calls can significantly increase token usage.
Retrieval Augmented Generation (RAG)
- Definition: Making large amounts of data available to AI agents by retrieving relevant information from a vector database.
- Process:
- Documents are broken into chunks.
- Data is stored in a vector database.
- AI agent queries the vector database semantically.
- Relevant information is retrieved and used to respond.
- Example: An AI agent answers questions about shipping policies by querying a vector store containing the company's documentation.
- Advanced Concepts: Hybrid RAG, agentic RAG, multimodal RAG.
- Context Engineering: Requires careful engineering due to the dynamic context.
- Implementation Detail: All the results from the vector store are dumped straight into the tool message and sent to OpenAI.
Context Isolation
- Definition: Separating tasks into sub-agents to manage individual contexts.
- Motivation: Limited number of tools an AI agent can handle, and too much context can be overwhelming.
- Implementation: Using multi-agent systems where each sub-agent manages its own memory and context.
- Example: An AI agent team generates a newsletter, with separate sub-agents for research, writing, publishing, analytics, and subscriber management.
- Benefits: Prevents the main agent's context from being polluted by external data.
- Architecture: Multi-layered multi-agent teams can be created.
Summarizing Context (Context Compression)
- Definition: Reducing the amount of information passed to the AI model by summarizing external data.
- Process:
- A sub-workflow is created to handle external data (e.g., scraping a website).
- The sub-workflow summarizes the data using an LLM chain or a sub-agent.
- The summarized data is passed back to the main agent.
- Example: A sub-workflow scrapes a website, summarizes the content, and returns the summary to the main agent.
- Benefits: Isolates context, reduces token usage, and improves efficiency.
Context Routing and Staging
- Definition: Deterministic route for how it's going to handle the context of such a large amount of data.
- Application: Deep research tasks that require careful management of large amounts of data.
- Example: The deep research blueprint in NAN takes a form submission (research topic) and does full deep research on that topic.
- Process:
- Research topics are researched one by one.
- Learnings are compiled.
- A chain of thought model generates a report using the insights.
- Considerations: Can be a very long-running task (up to an hour or more).
Context Formatting
- Definition: Changing the format of the context to make it more manageable and LLM-friendly.
- Example: Converting HTML to Markdown before passing it to the AI model.
- Benefits: Reduces token usage, speeds up processing, and improves the quality of the results.
- Alternative: Using services like firecrawl.dev that return Markdown directly.
Context Trimming
- Definition: Reducing the amount of data going back to the agent by trimming the context.
- Example: Taking only the first thousand characters from a document.
- Benefits: Reduces costs and speeds up processing.
- Other Methods: Reducing the context window length for short-term memory, reducing the number of chunks returned from a vector database.
Context Engineering Issues
- Context Poisoning: A hallucination makes its way into the context, causing the LLM to reproduce incorrect information.
- Context Distraction: Too much information in the context window makes it difficult for the LLM to pick out the relevant information.
- Unnecessary/Contradictory Information: Negatively affects the outcome.
Conclusion
Context engineering is a crucial skill for building effective AI agents and automations. It involves managing the context window, using techniques like short-term memory, long-term memory, RAG, tool calling, context isolation, summarization, routing, formatting, and trimming. By carefully managing the context, developers can mitigate issues like context poisoning and distraction, improve the efficiency of their agents, and achieve better results.
AI summaries can miss context or contain errors. Check important details against the original video.