THE SUMMARYAI-generated
Key Concepts
- Retrieval Augmented Generation (RAG)
- Multi-Agent Systems
- Long Context Models (e.g., GPT-4.1, Gemini)
- Index-Free Retrieval
- Chunking Strategy
- Scratchpad
- Content Routing
- Recursive Decomposition
- Answer Verification
- Workhorse Models
- Knowledge Graphs
1. Introduction to OpenAI's Multi-Agent RAG System
- The video addresses the common questions about RAG systems, specifically chunking strategies, embedding models, and vector stores.
- OpenAI has introduced a new multi-agent RAG system that mimics human information processing.
- This system is index-free, eliminating the need for chunking strategy and embedding model selection.
- It leverages long context models like GPT-4.1 (1 million tokens).
- The system combines multi-agent architecture with long context windows for improved retrieval results.
- The system is not universally applicable and is best suited for specific use cases.
2. Overview of the Approach
- The approach involves skimming the document to identify relevant sections.
- The document is divided into manageable subsections (e.g., chapters).
- The system identifies which sections are relevant to the user query and discards irrelevant ones.
- This process is repeated at multiple depths (a user-controlled hyperparameter).
- The final result is a set of subsections (paragraphs) used to generate the answer.
- A final verification step ensures the generated answers are grounded in the selected text, preventing hallucinations.
- Each step is treated as a sub-agent, requiring an LLM with specific capabilities.
3. Detailed Architecture and Implementation
- Initial Chunking: The document is divided into 20 chunks of equal size, ensuring each chunk ends with a valid sentence.
- First-Pass Evaluation: A long context LLM (e.g., GPT-4.1 Mini, Gemini Flash) evaluates each chunk along with the user question to determine relevance. The goal is to use a long context but inexpensive model.
- Scratchpad: The system uses a "scratchpad" to enable non-reasoning models to reason through the evaluation process.
- Multi-Step Process: Relevant sub-chunks are further broken down to identify specific paragraphs for accurate answers.
- Answer Generation: A more powerful model (e.g., GPT-4.1) generates structured output.
- Verification: A reasoning model (e.g., GPT-4 Mini) acts as a judge to validate the generated answer.
4. Use Cases and Limitations
- The system processes all text for each query without pre-created indexes.
- It is highly accurate but not suitable for every application.
- It is not ideal for simple document chats due to cost and latency.
- It is well-suited for report generation and factual data retrieval where latency is not critical.
- An example use case is complex question answering on long legal documents.
5. Practical Example: Trademark Trial and Appeal Board Manual
- The example uses the Trademark Trial and Appeal Board Manual of Procedure (1,200 pages).
- The manual is used as a reference in the patents office.
- Simply putting the entire document in the context of a long context model is not effective, especially for reasoning-intensive tasks.
6. Code Walkthrough and Implementation Details
- The video demonstrates a Google Colab notebook based on OpenAI's implementation.
- Document Loading: The document is downloaded and limited to 920 pages. Documents longer than the model's context window are divided into subdocuments and processed in parallel.
- Chunking: The document is split into chunks of equal size (approximately 33,000 to 62,000 tokens). The text is tokenized, and sentences are iteratively added to create chunks.
- Content Routing: The system uses a function to determine whether a chunk can answer the user's question.
- System Prompt: The system prompt instructs the LLM to act as an expert document navigator, record reasoning in a scratchpad, and select relevant chunks.
- "You are an expert document navigator. Your task is to identify which text chunk might contain information to answer the user question. Record your reasoning in a scratch pad for later reference. Choose chunks that are most likely relevant. Be selective but thorough. Choose as many chunks as you need to answer the question and first think carefully about what information would help answer the question then evaluate each chunk."
- Scratchpad Implementation: The scratchpad is a text string where the LLM adds its reasoning for chunk relevance. The LLM makes multiple calls to analyze chunks individually and select specific chunks.
- Recursive Decomposition: The function is called recursively to divide the document into smaller subsections until a specific depth is reached.
- Question Example: "What format should a motion to comply discovery be filed in? How should signatures be handled?"
- The system splits the document into 20 chunks, reasons through each chunk, and selects the most relevant ones. This process is repeated at multiple depths.
- Answer Generation: A system prompt is used to guide the answer generation process.
- "You are a legal research assistant answering questions about the trademark trials and appeal board manual of procedure. Answer questions based only on the provided paragraphs. Don't rely on any information or foundational knowledge or external information phrases of the paragraphs that are relevant to the answer."
- The system generates an answer with specific citations to relevant paragraphs.
- Answer Verification: An LLM acts as a judge to verify the answer's factual accuracy.
- "You are a fact checker for the legal information. Your job is to verify if the provided answer is factually accurate according to the source paragraphs. Use citations correctly. Be critical and look for any factual errors or unsupported claims. Assign a confidence level based on how directly the paragraph answers the question."
- The system assigns a confidence score to each paragraph.
7. Cost Analysis
- The agentic system has zero pre-processing cost but higher per-query costs due to multiple LLM calls.
- A hypothetical query costs approximately 36 cents, which is more expensive than a standard RAG system (40 cents for embedding and metadata generation, plus storage).
- The higher cost is justified in applications requiring very accurate answers, such as the legal domain.
8. Benefits and Trade-offs
- Benefits:
- Zero ingest latency
- Dynamic navigation
- Mimics human reading patterns
- Focus on promising sections
- Cross-section reasoning
- Trade-offs:
- Higher cost per query
- Increased latency
- Limited scalability in its current form
9. Potential Improvements
- Caching: Caching content can improve latency and reduce costs.
- Knowledge Graphs: Generating knowledge graphs can preserve relationships between entities.
- Scratchpad Improvement: Adding the ability to remove or edit information in the scratchpad.
- Depth Adjustment: Adjusting the depth of recursive decomposition based on the specific use case.
10. Conclusion
- The multi-agent RAG system offers high accuracy but comes with increased cost and latency.
- It is well-suited for specific use cases, such as legal document analysis, where accuracy is paramount.
- Future implementations may involve hybrid approaches combining traditional RAG with long context LLMs.
- The system leverages large context windows to achieve its capabilities.
- The video encourages discussion and feedback on potential improvements and implementations.
AI summaries can miss context or contain errors. Check important details against the original video.





