You NEED to Pick the Right RAG Pattern for Your Use Case (n8n)
By The AI Automators
Key Concepts
- RAG (Retrieval-Augmented Generation): A framework that combines retrieval of relevant information from a knowledge base with the generative capabilities of Large Language Models (LLMs) to produce more accurate and contextually relevant responses.
- Model Intelligence vs. Speed and Cost: The fundamental trade-off in RAG system design, where larger, more intelligent models offer higher quality but at increased computational cost and slower inference times, while smaller models are faster and cheaper but less capable.
- RAG Design Patterns: Nine distinct strategies for building RAG systems, categorized into deterministic and agentic approaches, each tailored to specific use cases and priorities.
- Deterministic RAG: RAG systems with pre-defined, step-by-step logic that execute the same way every time, ensuring reliability and predictability.
- Agentic RAG: RAG systems that utilize AI agents with tools and reasoning capabilities to dynamically decide on actions, retrieve information, and generate responses.
- Query Transformation: Techniques used to refine and expand user queries to improve retrieval accuracy, such as decomposition, rewriting, and expansion.
- Vector Store: A database optimized for storing and searching vector embeddings of text, enabling semantic similarity searches.
- LLM (Large Language Model): A type of artificial intelligence model trained on vast amounts of text data, capable of understanding and generating human-like text.
- Parameters: A measure of the size and complexity of an LLM, generally correlating with its capabilities.
- Context Window: The amount of text an LLM can process and consider at any given time.
- Tool Calling: The ability of an LLM to invoke external functions or APIs to gather information or perform actions.
- Multi-Agent Systems: RAG architectures where multiple AI agents collaborate to achieve a common goal, often with specialized roles.
- Sub-Agents: Specialized agents within a multi-agent system that handle specific tasks or domains.
- Sequential Chaining: A RAG pattern where agents or LLM calls are executed in a specific order, with the output of one step feeding into the next.
- Query Routing: Directing user queries to the most appropriate agent or processing path based on their intent or complexity.
RAG System Design: Balancing Intelligence, Speed, and Cost
The effectiveness of a RAG system is heavily dependent on its design, which must be tailored to the specific use case. A critical aspect of this design is striking the right balance between model intelligence (quality of responses) versus speed and cost. For instance, a customer-facing chatbot demands lightning-fast responses, making a slow, frontier model like Claude Sonnet impractical. Conversely, a RAG agent in a legal department prioritizes accuracy over speed, justifying the use of more complex models and reasoning. Resource constraints also play a significant role, particularly in fully local RAG systems where model selection is limited by available hardware. This variability in requirements leads to a multitude of RAG strategies, which the speaker has distilled into nine key design patterns.
Model Selection and its Impact
The choice of LLM significantly influences RAG system design. Generally, larger models with more parameters offer higher quality responses due to their superior instruction following, tool calling, complex reasoning, and context handling abilities. Small language models (1-10 billion parameters) often struggle with reliable tool calling and basic instruction following. A notable improvement occurs around the 15 billion parameter mark, with capabilities increasing up to frontier models like Claude Sonnet and GPT-5.
However, the speaker emphasizes that even small models (15-20 billion parameters) can outperform larger ones if they are provided with high-quality retrieval. This means that effective RAG system design, which injects the right content into the LLM's context, is paramount, regardless of model size.
RAG Use Cases and Design Priorities
The speaker outlines four distinct RAG use cases, highlighting how their unique priorities dictate system design:
1. Customer-Facing RAG Chatbot
- Priority: Speed is paramount to retain customer engagement.
- Model Choice: Smaller, faster models like Gemini 2.5 Flash are preferred for quick responses and lower operational costs.
- Scale: Deployed at scale to handle thousands of concurrent users, necessitating cost-effective solutions.
- Techniques: Likely to employ a mixture of agentic and non-agentic RAG, with guardrails for specificity, query routing for directing questions to appropriate tools, and the verify answer pattern to ensure factual accuracy.
- Constraint: Avoids expensive frontier models ($15 per million output tokens).
2. AI Assistant/Co-pilot (e.g., Legal Department AI Agent)
- Priority: Accuracy is crucial, especially with complex documents.
- Model Choice: Larger, more capable models are used to facilitate complex reasoning.
- Trade-off: Speed is reduced due to the need for in-depth document analysis and potentially iterative/recursive retrieval.
- Cost: Higher per-user cost due to more expensive models and larger context windows, but acceptable for a smaller, internal user base.
- Deployment: Can be integrated via tools like N8N's MCP trigger, Open Web UI, or as Slack/Microsoft Teams agents.
- Techniques: Agentic RAG, multi-agent RAG, hybrid RAG, and the verify answer pattern.
3. AI Automation with Embedded RAG
- Example: An agentic blogging system that retrieves information, generates article outlines, and then builds articles.
- Priority: Resilience and background operation are key; speed is less critical.
- Model Choice: Larger models with reasoning capabilities are suitable.
- Techniques: Agentic RAG, non-agentic RAG, deterministic flows, and extensive prompt chaining.
4. Fully Local RAG System
- Constraint: Model selection is severely limited by available internal hardware (e.g., a $3,000 graphics card might only support a 20 billion parameter model).
- Design Influence: Hardware limitations heavily influence system design.
- Cost: High initial capital cost for hardware, but low long-term cost per user.
- Inference Speed: Dependent on the quality of graphics cards.
- Deployment: Often uses interfaces like Open Web UI to maintain local operation.
Deterministic RAG System Designs
These patterns focus on pre-defined, step-by-step execution for reliability.
Pattern 1: Naive RAG
- Description: The most basic RAG system. A user message is directly sent to a vector store for retrieval, then to an LLM for response generation, and finally outputted.
- Pros: Ultra-fast and simple.
- Cons: Highly basic and prone to errors. The LLM has a single pass to generate an answer, and if the retrieved chunks are irrelevant, the answer will be incorrect or "I don't know." Conversational queries with stop words can negatively impact embedding quality.
- Example: A query about changing power levels on a GE Advantium oven, when only chunks related to a different GE model (Cafe) are retrieved, results in an incorrect "I don't know" response.
- Implementation (N8N): Chat trigger -> Vector Store -> LLM call.
- Performance: End-to-end took 2.5 seconds, including vector search and LLM inference.
Pattern 2: Deterministic RAG with Query Transformation and Verify Answer
- Description: Enhances Naive RAG by incorporating query transformation techniques and a verification step. This pattern does not involve tool calling.
- Process:
- Query Intent and Decomposition: The incoming query is broken down into sub-queries.
- Retrieval Gate: Checks if retrieval is necessary (e.g., not for small talk).
- Query Rewriting and Expansion: Sub-queries are rewritten with synonyms, related terms, and different phrasing to broaden the search.
- Product Category Classification: The LLM classifies the query into a product category to narrow down the vector store search using metadata filters.
- Query Clarification (if needed): If the product category is unclear, the system asks the user for clarification.
- Query Routing: Queries are directed to the most appropriate metadata filter (e.g., "oven").
- Multiple Vector Searches: Expanded queries are used to perform multiple searches against the vector store, with metadata filters applied.
- RAG Fusion (Reciprocal Rank Fusion): Results from multiple searches are fused and de-duplicated into an ordered list. Chunks appearing in multiple lists are ranked higher.
- Re-ranking: A cross-encoder model (similar to a basic LLM) re-ranks the top results based on the user's question and retrieved context, selecting the most relevant ones.
- Answer Generation: The LLM generates an answer using the user's question, chat history, and the top re-ranked chunks.
- Verify Answer Pattern: The generated answer is checked against the retrieved context for accuracy, grounding, and contradictions. If inaccurate, it's sent back to the LLM for revision.
- Output: The verified answer is presented to the user, and memory is updated.
- Pros: Significantly improves accuracy and relevance compared to Naive RAG. Can be run on smaller LLMs due to the breakdown of complex tasks into multiple LLM calls.
- Example: A multifaceted question about a non-existent "GE Abacus mixer" is decomposed, clarified, routed to "ovens," and then processed through multiple retrieval and verification steps to provide an accurate response.
- Performance: Took 30 seconds end-to-end, with the answer verification pattern being a significant bottleneck (9 seconds).
- Key Feature: Deterministic and highly reliable.
Pattern 3: Deterministic RAG using Iterative Retrieval
- Description: An extension of the previous pattern where, after re-ranking but before answer generation, an LLM analyzes the retrieved results. If the quality is insufficient, it triggers further retrieval.
- Process:
- ... (Steps similar to Pattern 2 up to re-ranking)
- Analyze Results (LLM Call): Assesses the quality and sufficiency of retrieved chunks.
- Iterative Retrieval Decision: Determines if more retrieval is needed.
- Iteration Counter: Tracks the number of retrieval iterations to prevent infinite loops.
- Further Retrieval (if needed): If more retrieval is required, new queries are generated and sent back to the vector store.
- Answer Generation: Once sufficient information is retrieved or the iteration limit is reached, the LLM generates the final answer.
- Pros: Achieves very high-quality responses by iteratively refining the retrieval process.
- Cons: Increases response time.
- Implementation (N8N): Involves an "Analyze Results" LLM call, an "if" node for retrieval decision, and a counter to manage iterations.
Pattern 4: Adaptive Retrieval
- Description: A more sophisticated deterministic pattern that dynamically adjusts the retrieval strategy based on query complexity.
- Process:
- Query Classifier: Determines if retrieval is needed at all.
- No Retrieval: The LLM uses its training data for a simple response.
- Retrieval Required: Proceeds to determine the retrieval strategy.
- Retrieval Strategy Determination:
- Single-Step Retrieval: A single query (or expanded query) is sent to the vector store.
- Multi-Step Retrieval: A recursive pattern is initiated.
- First query to the vector store, results are retrieved and re-ranked.
- Analyze Results Node: Assesses if retrieval is complete or if more digging is needed for complex multi-step queries.
- Generate More Queries: If necessary, new queries are generated and sent back to the vector store.
- This cycle continues until retrieval is complete.
- Answer Generation: Once retrieval is finished, the LLM generates the answer.
- Query Classifier: Determines if retrieval is needed at all.
- Analogy: Similar to how GPT-4 routes simple questions to smaller models and complex ones to larger, more reasoning-intensive models.
- Implementation (N8N): Uses a "Query Intent" node for classification and a retrieval decision, followed by "if" nodes to branch into single-step or multi-step retrieval, with multi-step involving an "Analyze Results" node and loops back to the vector store.
Synthesis of Deterministic Patterns: These patterns offer high reliability and predictability. They are ideal when consistent, repeatable results are crucial. While they can be implemented with smaller LLMs, the complexity is managed through structured workflows and multiple LLM calls rather than relying solely on the LLM's inherent intelligence. The combination of deterministic workflows with agentic capabilities (hybrid approach) is presented as a powerful strategy for production RAG.
Agentic RAG Patterns
These patterns leverage AI agents for more dynamic and flexible RAG systems.
Pattern 5: Standard Agentic RAG
- Description: An AI agent receives a user message and uses a set of tools (e.g., vector search, database search, web search) to retrieve context and generate a response.
- Key Features:
- Tools: Agents can call various external services.
- System Prompt: Complex system prompts can guide the agent's behavior, including standard operating procedures and prompt engineering techniques.
- Memory: Agents can retain conversation history for context.
- Function Calling Loop: Enables agents to make multiple tool calls, including different variations of a query, effectively performing query transformation implicitly.
- Pros: Highly flexible, can implicitly handle query transformation and decomposition due to its reasoning and tool-calling capabilities. Frontier models excel here.
- Cons: Smaller LLMs (under 10 billion parameters) may struggle with reliable tool calling. The quality of responses heavily depends on the LLM's intelligence and the system prompt.
- Example: An agent can automatically refine queries like "GE Abacus mixer" to "Advantium oven power levels" and "cleaning instructions" by hitting the vector store multiple times with different keywords.
- Implementation (N8N): Involves an AI agent node with tools, memory, and an LLM.
- Model Comparison: Claude Sonnet 4.5 (frontier model) reliably identifies non-existent products and asks for clarification. GPT-OSS 20 billion (smaller model) can sometimes fabricate information when asked repeatedly, demonstrating less reliability.
Pattern 6: Hybrid RAG
- Description: Similar to Standard Agentic RAG, but the agent can retrieve information from diverse data stores beyond just vector databases.
- Additional Tools: Includes tools for database search (e.g., writing SQL queries for PostgreSQL) and graph search (e.g., writing Cypher queries for Neo4j).
- Pros: Enables retrieval from structured and semi-structured data sources, providing a more comprehensive knowledge base.
- Cons: Requires agents to be proficient in generating queries for different data types.
- Related Content: The speaker references videos on "Agentic Databases" and "Graph Agents" for further details.
Pattern 7: Multi-Agent RAG System using Sub-Agents
- Description: A system where a main orchestrator agent delegates tasks to specialized sub-agents. This is an alternative to having a single, highly complex system prompt for a powerful LLM.
- Benefits:
- Simplified System Prompts: The main agent's prompt can be simplified, with responsibilities offloaded to specialized sub-agents.
- Context Window Management: Sub-agents can handle tasks that require large context windows (e.g., summarizing an entire document) without overwhelming the main agent's context. The sub-agent performs the task and returns a summary, allowing the main agent to continue with its workflow without needing the full document in its context.
- Specialization: Sub-agents can be highly specialized for specific tasks, like database queries or document research.
- Example: A "database sub-agent" with numerous SQL tool calls, or a "document researcher" that fetches and summarizes a document.
- Implementation (N8N): Involves creating AI agent tools that act as sub-agents, which can then be called by the main agent.
- Caution: Requires highly specific roles and responsibilities for each agent. Building overly complex multi-agent systems (e.g., 25 sub-agents) can become unwieldy and unreliable. A conservative approach, starting with one sub-agent and gradually adding more, is recommended.
Pattern 8: Multi-Agent RAG with Sequential Chaining
- Description: Agents or LLM calls are executed in a sequential flow, where the output of one step feeds into the next. This is more aligned with deterministic workflows but uses agents.
- Process:
- Input: A chat message or an automated trigger.
- Agent 1: Performs a task (e.g., research using various tools).
- Agent 2: Receives the output of Agent 1 and performs another task (e.g., writing a blog post).
- Agent 3: Receives the output of Agent 2 and performs a final task (e.g., publishing to WordPress).
- Pros: Allows for complex automations with specialized agents at each stage. Can use simpler prompts for each agent.
- Example: An AI blogging system where a researcher agent gathers information, a writer agent drafts the content, and a publication agent posts it. This can also include human-in-the-loop steps.
- Flexibility: Can incorporate non-agent LLM calls for text generation or JSON output.
Pattern 9: Multi-Agent RAG System with Routing and Sequential Chaining
- Description: Combines query routing, sequential chaining, and multi-agent capabilities. Queries are first classified and routed to the most appropriate agent or workflow.
- Process:
- Query Classifier and Router: Analyzes the user's query to determine the best path.
- Routing:
- Simple Questions: Directly to a simple response generator.
- Writing Tasks: To a dedicated writing agent.
- Research Tasks: To a dedicated research agent.
- Sequential Chaining (if applicable): If a task involves multiple steps (e.g., research and drafting), it can be routed to a combined research and writing agent or a sequence of agents.
- Pros: Speeds up inference by directly routing queries to specialized agents, avoiding unnecessary orchestration overhead. Maintains separation of concerns and specialized agents with simpler prompts.
- Implementation (N8N): Uses a query classifier and router node to direct the flow to different agentic or deterministic paths.
Other Mentioned RAG Patterns and Techniques
- Selfrag: Uses reflection tokens to assess retrieval quality and trigger re-retrieval.
- Corrective RAG: Falls back to web search if content is not found in the knowledge base.
- Guardrails: Protects against prompt injection attacks and Personally Identifiable Information (PII) leakage.
- Human in the Loop: Escalates to human interaction for complex cases or automations.
- Human Handoff: Directs chat messages to a human support agent when the AI cannot answer.
- Multi-step Flows: Traditional conversational chat flows where sequential questions build upon each other.
- Context Expansion: A technique to load neighboring chunks, document sections, or the entire document into context.
- Deep RAG: An emerging trend involving "deep agents" and retrieval planning for longer-term execution.
- Lexical Keyword Search, Dynamic Hybrid Search: A specific technique not fully detailed but highlighted for its importance.
Conclusion
The speaker emphasizes that there is no single "best" RAG system. The optimal design is highly use-case dependent, requiring a careful consideration of the trade-offs between model intelligence, speed, and cost. The nine design patterns presented offer a comprehensive framework for building robust and effective RAG systems, ranging from simple deterministic flows to complex multi-agent architectures. The ability to implement these patterns in tools like N8N empowers developers to create production-ready RAG applications. The key takeaway is that understanding these patterns and their underlying principles is crucial for RAG project success.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

LIVE: Pope Leo AI-focused encyclical
Reuters

Bring back borstals! Top cop slams parents lack of discipline | The Daily T
The Telegraph

BREAKING: Huge US-Iran News Dropping Today!
The Economic Ninja

College Kids Don’t Want Your AI
Bloomberg Television

SpaceX Starship Successfully Deploys Mock Satellites
Bloomberg Television

‘Very encouraging’: Pauline Hanson on shock new seat-by-seat modelling showing One Nation wins
Sky News Australia

Top five wild lefty lunacy moments in the United States
Sky News Australia