Key Concepts
Retrieval Augmented Generation (RAG), Agentic RAG, Vector Database, Embedding Model, Local AI, n8n, Ollama, Context Length, PostGres, Superbase, Document Chunks, Agent Tools, Knowledge Base, SQL Queries.
Local Agentic RAG with n8n: A Comprehensive Guide
Introduction
This video introduces a local version of an n8n agentic RAG template, enabling users to create a secure, offline AI agent. It builds upon the cloud version previously showcased, addressing user requests for a local AI implementation and emphasizes the importance of controlling data and LLMs locally.
What is Agentic RAG?
- Naive/Traditional RAG: Documents are split into chunks, converted to vectors using an embedding model, and stored in a vector database. User questions are similarly converted into vectors and matched against document vectors. The resulting chunks, along with the question and system prompt, are fed to the LLM. This is a "one-shot" process. The LLM doesn't have the chance to refine the query, choose a different vector database, or explore the data differently.
- Agentic RAG: Introduces an agent before retrieval. The agent has tools to perform RAG and explore the knowledge base in different ways. This allows for reasoning during the retrieval process, querying multiple vector databases, and using various knowledge exploration methods.
"Instead of there being a retrieval before we reach the llm with agentic Greg, the agent comes first and then it has tools available for it to perform Rag and then also this is a very key thing explore the knowledge base in other ways"
n8n Agentic RAG Template: Local Implementation
The local n8n template mirrors the cloud version, replacing cloud services with local alternatives like Ollama and self-hosted Superbase. The agent is equipped with tools to interact with Postgres (or Superbase), enabling diverse knowledge base exploration methods.
Use Case Examples and Testing
- Basic RAG Retrieval: Asking "Who is the CEO of the company?" triggers a simple RAG lookup, successfully identifying Dr. Alisa vanderhoven.
- Complex Question Requiring Document Listing: Asking "Who was in the Feb 12th meeting?" prompts the agent to list available documents and then read the relevant meeting minutes to extract the attendee list.
- Tabular Data Analysis: Asking "What is the average overall score based on the customer feedback?" triggers the execution of a SQL query against a customer feedback table, demonstrating the agent's ability to perform calculations on tabular data (average overall rating of 7.89).
The presenter notes that results might vary depending on the local LLM used, with smaller models producing less consistent outcomes.
Workflow Setup and Explanation
-
Database Setup: Utilizing PostGres for storing conversation history and the knowledge base. The workflow includes nodes to create necessary tables:
document_metadata: Stores high-level document information, enabling the LLM to list available documents.document_rows: Stores tabular data from Excel/CSV files in a JSONB column (row_data), accommodating different table structures within a single table.documents: Stores document embeddings for RAG.
-
RAG Pipeline (Blue Box): Adds local files to the Superbase/PostGres knowledge base.
- Uses a local file trigger to watch for added or changed files in a specified folder.
- Deletes existing document records to ensure the knowledge base contains only updated information.
- Inserts document metadata (ID, title, creation date, schema).
- Reads the file content and extracts text based on file type (PDF, CSV, TXT).
- For text files: Uses the default data loader to extract text and create vector embeddings.
- For CSV/Excel files: Extracts data, inserts rows into the
document_rowstable, and aggregates data for regular RAG.
- Sets the schema for tabular data in the
document_metadatatable, allowing the agent to understand the structure and write SQL queries.
-
AI Agent Setup (Yellow and Green Boxes):
- Uses chat and webhook triggers for interaction.
- Employs a basic system prompt defining the agent's role and available tools.
- Utilizes PostGres chat memory for storing conversation history.
- Leverages the Open AI node (pointing to a local Ollama instance) as a workaround due to a bug with the Chat Ollama Model node.
- Defines tools for RAG (PostGres PG Vector), listing documents, retrieving file contents, and querying tabular data.
Addressing the Context Length Limitation with Ollama
Ollama models have a default context length of 2,000 tokens, which can lead to issues with longer prompts and tool invocations.
- Problem: The agent's system prompt and tool instructions can be truncated due to the limited context window.
- Solution: Create a custom model file inheriting from the base model and modifying the context length parameter.
- Example command:
ollama create <model_name> -f <model_file> - Modify the
MODEL_FILEto include:
- Example command:
FROM <base_model>
PARAMETER CONTEXT_LENGTH 8192
After modifying the MODEL_FILE, you can create a new model from the model file, which will be based on the original model but using the parameters that are specified in the model file.
This allows the agent to handle more complex queries and retain context effectively, as it will be able to consider prompts up to the maximum token number you specified in the MODEL_FILE.
Cartesia Sponsorship
The video is sponsored by Cartesia, a company developing voice AI technology for voice agents, cloning, changing, and text to speech. Their Sonic 2.0 model is based on a state space model architecture and delivers low latency and control over emotions, accents, and speed. The playground allows voice cloning with as little as 3 seconds of audio.
"gentic rag with local AI is an absolute Game Changer like that definitely sounds like me I am very impressed"
Synthesis/Conclusion
The local n8n agentic RAG template provides a powerful, secure, and offline solution for building AI agents. By leveraging local LLMs and databases, users maintain control over their data and can create domain experts on their documents. Overcoming the limitations of naive RAG with agentic techniques allows for more flexible and intelligent knowledge base exploration. Addressing the context length limitation of Ollama ensures reliable performance with complex queries. This setup can be further expanded for specific use cases, showcasing the power of local AI and no-code tools.
AI summaries can miss context or contain errors. Check important details against the original video.