This RAG Trick Makes Your AI Agents WAY More Accurate (n8n)

By The AI Automators

Share:

Key Concepts

  • RAG (Retrieval Augmented Generation) Agents: AI agents that retrieve information from a knowledge base before generating a response.
  • Vector Store: A database that stores numerical representations (embeddings) of text chunks, enabling semantic search.
  • Hybrid Search: Combines vector search (semantic) with keyword search for improved retrieval.
  • Context Expansion: The ability for a RAG agent to intelligently retrieve additional, related sections or parts of a document to provide comprehensive context.
  • Neighbor Expansion: Fetching chunks immediately adjacent (before and after) a retrieved candidate chunk.
  • Parent Expansion: Retrieving all chunks belonging to the parent section or heading of a retrieved chunk.
  • Agentic Expansion: The most sophisticated approach, where an AI agent uses a document's inherent hierarchy to fetch multiple relevant sections or even the entire document.
  • Recursive Character Text Splitter: A standard chunking method that splits text based on character count and delimiters like newlines.
  • Markdown Splitting: A chunking method that prioritizes markdown headings (H1, H2, etc.) as split points.
  • Contextual Embeddings: An approach where an LLM generates a one-sentence snippet explaining the context of a chunk, added to its metadata.
  • Query Expansion: The agent sends multiple, varied queries to the vector store to broaden the search.
  • LLM Chain: A sequence of LLM calls and other operations, often more deterministic than an agent.
  • N8N: A workflow automation tool used for implementing the RAG system.
  • Superbase: A backend-as-a-service platform, used here for its Postgres database capabilities, serving as both vector store and record manager.
  • Postgres: A relational database, recommended for RAG systems due to its ability to perform both SQL queries and vector searches.
  • OCR (Optical Character Recognition): Used to extract text, including headings, from documents like PDFs.
  • Record Manager: A system (often a database table) to track documents upserted into a vector store, preventing duplication and managing metadata.
  • Hierarchical Index: A structured representation of a document's headings and their corresponding chunk ranges, enabling navigation.

Introduction: The Fundamental Problem with RAG Agents – Lost Context

The primary reason RAG agents often fail is their inability to grasp the "big picture." They generate responses from isolated fragments (chunks) of a document, completely overlooking the document's inherent structure that provides meaning to these fragments. For instance, an agent might retrieve a chunk stating a policy update occurred last month but remains unaware of the policy's content, the specific changes, or their impact. This loss of structural meaning during document splitting and chunking is a significant problem, frequently leading to hallucinations.

A key example provided is an insurance policy knowledge base. If asked, "Is tennis elbow covered under this policy?", the agent might retrieve fragments discussing "coverage" and "tennis elbow." However, if the "tennis elbow" fragment is from a "policy exclusion" section, the agent, lacking this structural context, might incorrectly assume it's covered. This is a faithfulness problem, not necessarily hallucination, but a failure to receive sufficiently accurate information from the vector store.

Existing mitigation strategies have limitations:

  • Contextual Embeddings: Involves an LLM creating a one-sentence snippet for each chunk (e.g., "This is in the policy exclusion section"). This is costly and not scalable due to an LLM call per chunk.
  • Query Expansion: The agent sends multiple queries from different angles, but this approach lacks reliability.

The Solution: Context Expansion

Context expansion is presented as a superior and scalable solution. It enables the agent to intelligently retrieve entire sections, subsections, or related parts of a document, providing all necessary information for a comprehensive and accurate answer without requiring an LLM call for every chunk.

Approaches to Context Expansion

The video details five methods for context expansion, implemented in N8N:

1. Full Document Expansion

  • Process: The user's question triggers a vector store search for candidate chunks. The AI agent identifies a "golden chunk" containing the relevant doc ID. It then uses this doc ID to fetch the entire document into context.
  • Pros: Provides the most comprehensive and accurate context, ensuring the agent has all information. Ideal for small documents (e.g., 3 pages).
  • Cons: Can be very expensive per query for large documents (e.g., 200 pages) due to increased token usage.
  • N8N Implementation:
    • Requires doc ID to be stored in the chunk's metadata during ingestion.
    • Leverages Superbase (built on Postgres) for its relational database capabilities. An SQL SELECT statement retrieves all content and metadata from the documents table where metadata->>'doc ID' matches the identified ID.
    • The retrieved chunks can be ordered by their internal ID to maintain the original document sequence.
    • Optimization: The total character count of a document can be stored in metadata. The agent can be instructed to trigger full document expansion only if the document size is below a certain threshold, managing cost and latency.
    • Deterministic LLM Chain: This approach can be implemented using a more deterministic LLM chain (rewriting query, reranking, deciding best chunk, if node for size check, Postgres query, final LLM for response) for potentially faster and cheaper execution with smaller models, compared to a reasoning agent.

2. Neighbor Expansion

  • Process: After retrieving a candidate chunk, the agent triggers a tool call to fetch the chunk immediately before and after the selected chunk.
  • Pros: Relatively simple to implement.
  • Cons: The agent still "doesn't know what it doesn't know." It retrieves arbitrary adjacent chunks without understanding their structural relevance, potentially still missing crucial context. This method can be imperfect if there are many consecutive newlines in the source document.
  • N8N Implementation:
    • Relies on line numbers stored in chunk metadata (standard in N8N's vector store integrations).
    • An SQL query searches for chunks with the same doc ID and line numbers that are decremented by one (for the previous chunk) or incremented from the end of the current chunk (for the next chunk).

3. Section and Parent Expansion (Leveraging Document Structure)

  • Goal: To fetch all chunks belonging to a specific section or under a parent heading, providing a more focused and structurally relevant context than neighbor expansion.
  • Example: For the "tennis elbow" question, this would retrieve the entire "Policy Exclusion" section, clearly indicating it's not covered.
  • Reliance on Document Structure: This approach is highly effective for structured documents common in enterprises (policies, regulations, reports) that have inherent markdown-like headings (H1, H2, H3).

The Gold Standard of Chunking: Custom Markdown Splitting

  • Problem with Standard Chunking:
    • Recursive Character Text Splitter: Chunks often span multiple, unrelated topics (e.g., "dimensions" and "suspension" in one chunk), leading to less focused embeddings and reduced retrieval accuracy.
    • Native N8N Markdown Splitting: While better, it still uses a chunk size and tracks back for headings, meaning a chunk can still start in one section and encompass parts of another if the chunk size is large.
  • Optimal Approach:
    1. First Pass: Split by Markdown Headings: The document is initially split into logical sections based on H2, H3, H4 headings, effectively creating "subdocuments." This ensures sections like "suspension fairings" are isolated from "dimensions."
    2. Second Pass: Recursive Character Text Splitting within Sections: For large sections created in the first pass, recursive character text splitting is then applied to break them into smaller, manageable chunks. This ensures chunks are topic-focused and within their correct structural boundaries.
  • N8N Implementation: This two-pass chunking is not natively supported by N8N's default data loader. It requires custom code nodes to implement this logic.
  • Demo (Impava Oven): The agent retrieves a candidate chunk. The custom chunking process enriches the chunk's metadata with child range (chunk indexes for its section) and parent range (chunk indexes for its parent section). The agent uses these ranges to pass to a Superbase Edge Function, which then retrieves all chunks within the specified range(s).
  • Rich Metadata: The custom chunking provides highly detailed metadata, including a cascading path (e.g., H1: "Using your washer", H2: "Begin procedure"), and chunk ID/index ranges for both the current section and its parent.
  • Metadata Enrichment (LLM per Document): An LLM is used once per document (not per chunk, making it scalable) to extract high-level metadata like brand, appliance type, and a brief document summary. This is injected into the metadata.
  • Contextual Snippets: By combining the document summary and the section heading from the rich metadata, a "contextual snippet" can be automatically generated and prefixed to each chunk (e.g., "This chunk is from an EPA 24in single wall oven instruction manual. Specifically the installation section part two."). This achieves similar benefits to costly contextual embeddings without per-chunk LLM calls.

4. Agentic Expansion using Document Hierarchy (Most Sophisticated)

  • Process:
    1. The agent retrieves a "golden chunk" with its doc ID.
    2. Instead of immediate expansion, it fetches the entire document hierarchy (the mapping of headings to chunk ranges) for that document.
    3. Crucially, if the retrieved chunk references multiple parts of the document (e.g., "as discussed in section two," "definition section," "appendix"), the agent can use the hierarchy to identify and pass multiple specific chunk ranges to the context expansion endpoint.
    4. With all relevant chunks in context, it generates a comprehensive response.
  • Benefits: Allows the agent to "navigate" the document structure intelligently, pulling information from disparate but relevant sections without needing a complex knowledge graph.
  • Demo (Impava Oven): The agent uses a fetch document hierarchy tool. This hierarchy (showing H1, H2, H3 levels with their respective chunk ranges) informs the agent which ranges to pass to the Superbase context expansion endpoint.
  • Superbase Edge Function & Database Function: An Edge Function acts as an API endpoint, validating input and triggering a Postgres database function (get_chunks_by_ranges). This database function loops through an array of doc IDs and chunk index ranges to perform SQL SELECT queries and return the specified chunks.
  • Saving the Hierarchy: The document hierarchy (the hierarchical index) is saved in a record manager table in Superbase (Postgres) as a dedicated column. This is a key advantage of using Postgres, as pure vector stores like Pine Cone or Quadrant cannot store this complex structure directly.

N8N Implementation Details: The Ingestion Pipeline

The video outlines a sophisticated N8N ingestion pipeline for structured documents:

  1. File Trigger: Grabs new files (e.g., from Google Drive).
  2. Document Data Setup: Sets basic document data.
  3. OCR for Text Extraction: Uses Mistral OCR (because N8N's native PDF node doesn't extract headings) to extract text and crucially, headings.
  4. Page Number Aggregation: Aggregates page numbers for each chunk, enabling traceability (e.g., "chunk extracted from pages 11 and 12").
  5. Record Manager Interaction:
    • If a new document, creates a new row.
    • If an existing document with changed content, deletes previous vectors to prevent duplication and re-ingests.
    • Saves the hierarchical index (document structure and chunk ranges) in a dedicated column.
  6. Document & Metadata Enrichment (LLM Call): A single LLM call per document enriches metadata with brand, appliance type, and a document summary.
  7. Smart Markdown Chunker & Document Hierarchy Extractor (Custom Code Nodes): This is the core custom logic:
    • Parses by headings, then performs recursive character text splitting within sections.
    • Smartly merges very small chunks that might otherwise "pollute" the vector store.
    • Adds the contextual snippet (heading prefix) to each chunk.
    • Extracts the document hierarchy and maps sections to chunk ranges.
  8. Embedding Generation & Direct Injection into Superbase: Embeddings are generated, and chunks are injected directly into the Superbase vector store via a Postgres node. This direct injection is preferred over N8N's native Superbase vector store node for custom chunking at scale, as it performs better with chunk-specific metadata.

Conclusion

Context expansion is presented as a vital component for building accurate and comprehensive RAG agents, addressing the fundamental problem of lost context. The detailed N8N implementation, particularly the custom two-pass chunking, rich metadata enrichment, and agentic expansion leveraging a document hierarchy stored in Postgres, offers a robust and scalable solution. This approach allows RAG agents to "navigate" documents intelligently, providing highly relevant and complete answers. The system is being integrated into the AI Automators community's state-of-the-art RAG system, which also supports tabular data, knowledge graphs, diverse file formats, and dynamic hybrid search.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video