How to Infinitely Scale Your n8n RAG Workflows

The AI AutomatorsAbout 8 min readSep 17, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Orchestrator Workflow: A workflow that manages and parallelizes the execution of other workflows (e.g., the rag ingestion workflow).
  • Rag Ingestion Workflow: A workflow that processes files, chunks data, and stores it in a vector database.
  • Parallel Processing: Executing multiple instances of a workflow simultaneously to improve processing speed.
  • Error Handling: Implementing mechanisms to automatically retry failed executions and prevent workflow interruptions.
  • Q Mode: A feature in N8N that enables queuing and parallel execution of workflows.
  • Concurrency: The ability to run multiple tasks or workflows at the same time.
  • Webhooks: A way for applications to communicate with each other in real-time.
  • Subworkflows: Workflows that are called from within other workflows.
  • Vector Database: A database that stores data as vectors, which are numerical representations of data points.
  • Superbase: A backend-as-a-service platform that provides a vector database.
  • SFTP: Secure File Transfer Protocol, a secure way to transfer files between computers.
  • Upserts: A database operation that inserts a new row if it does not exist, or updates an existing row if it does.

Scaling Rag Ingestion with N8N: An Orchestration Approach

The Challenge of Scaling Rag Systems

The video addresses the challenges of scaling a Retrieval-Augmented Generation (RAG) system's data ingestion pipeline using N8N. While importing a few files is straightforward, scaling to thousands of documents can lead to:

  • Server Crashes: Due to memory overload or database overload.
  • Rate Limit Errors: From providers like OpenAI or Superbase.
  • Infrastructure Limits: Exceeding compute or storage capacity on N8N or external services.
  • Workflow Stalls: Due to unexpected errors and edge cases.
  • Long Processing Times: Processing thousands of files sequentially in a single workflow can take weeks.

The Orchestrator Workflow Solution

The video presents an orchestrator workflow as a solution to these challenges. The orchestrator workflow:

  • Parallelizes Ingestion: Calls the main rag ingestion workflow multiple times in parallel.
  • Handles Errors Gracefully: Automatically retries failed executions.
  • Manages Batch Processing: Processes files in batches to avoid overloading the system.

The goal is to reliably import tens of thousands of files with minimal manual intervention. The video focuses on building the orchestrator workflow and error handling, while the next video will cover scaling N8N itself and tuning the ingestion workflow for performance.

Orchestrator Workflow Demo

The demo showcases the orchestrator workflow in action:

  1. File Listing: The workflow retrieves a list of files from an SFTP server.
  2. Batching: The files are divided into batches of 50.
  3. Parallel Execution: The main rag ingestion workflow is executed 50 times concurrently for each batch.
  4. Status Tracking: A database table (Q table) is used to track the status of each execution.
  5. Error Retries: If an error occurs, the orchestrator retries the failed execution.
  6. Continuous Processing: After a batch is completed, the orchestrator processes the next batch until all files are ingested.

The demo highlights the speed and efficiency of parallel processing compared to sequential processing within a single workflow.

Orchestrator Workflow Architecture

The orchestrator workflow consists of the following components:

  1. File Listing (SFTP): Retrieves a list of files from an SFTP server.
    • Uses SFTP account credentials and the file path.
    • Filters out folders to avoid errors.
  2. Loop Over Items (Batching): Splits the file list into batches of 50.
    • Uses the "Loop Over Items" node (formerly "Split In Batches").
    • Batch size is set to 50.
  3. Superbase (Q Parent Executions): Creates a new parent execution row in the Q parent executions table for each file in the batch.
    • Stores the orchestrator execution ID, file name, and file path.
    • Returns the new ID created in the Superbase database (parent Q ID).
  4. Merge: Combines the parent execution IDs with the file data from the loop over items.
    • Uses "Combine by Positions" mode.
  5. Aggregate and Split Out: Used to accommodate error handling later on.
  6. Loop Over Items (Individual File Processing): Iterates over each file in the batch.
    • Batch size is set to 1.
    • Important: Uses the "Reset Condition" (node name and context done) to prevent issues with nested loops.
  7. Webhook Call (Rag Ingestion Workflow): Calls the rag ingestion workflow via a webhook for each file.
    • Passes the file data (ID, name, path, first attempt) to the ingestion workflow.
    • Creates a one-to-one mapping between a file and an execution of the rag ingestion workflow.
  8. Superbase (Child Executions): Creates a new child execution row in the Q child executions table for each execution.
    • Stores the child execution ID, parent Q ID, and status (initially "pending").
    • Uses a foreign key to enforce data integrity between the child and parent execution tables.
  9. Aggregate: Aggregates all the items in the batch into one item.
  10. Wait: Waits for 5 seconds before querying the database for updates.
  11. SQL Query (Check for Errors): Queries the database to check if any files have errored out.
    • Uses a common table expression (CTE) to count and aggregate data from the Q child executions table.
    • Checks if any files are not currently processing, have not been successful, and have not exceeded the maximum number of retry attempts.
  12. Router: Routes the workflow based on the results of the SQL query.
    • If errors are found, routes back to the loop over items to reprocess the failed files.
    • If no errors are found, proceeds to the next SQL query.
  13. SQL Query (Check for Pending Executions): Queries the database to check if any executions are still in progress.
    • Checks if the inflight count (number of pending executions) is greater than zero.
  14. Router: Routes the workflow based on the results of the SQL query.
    • If pending executions are found, loops back to the wait node to recheck later.
    • If no pending executions are found, loops back to the start to process the next batch.

Rag Ingestion Workflow Modifications

The main rag ingestion workflow requires some modifications to be triggered by the orchestrator:

  1. Webhook Trigger: Replace the original trigger (e.g., Google Drive trigger) with a webhook trigger.
  2. Data Input: Pass the relevant file data (ID, name, path, first attempt) via the webhook.
  3. Queue Update: Update the Q table at the end of the workflow to indicate the status of the execution (success or error).
  4. Error Handler Workflow: Create a separate error handler workflow that is triggered when an error occurs in the main ingestion workflow.

Error Handler Workflow

The error handler workflow:

  1. Error Trigger: Starts with an error trigger that is activated when an error occurs in the main ingestion workflow.
  2. Wait: Waits for 5 seconds to allow for timing differences in the creation of the row on the orchestrator side.
  3. Superbase (Update Row): Updates the Q child executions table to mark the execution as "error" and store the error message.
    • Uses the execution ID passed from the main ingestion workflow to identify the correct row to update.

Webhooks vs. Subworkflows

The video explains why webhooks are preferred over subworkflows for this use case:

  • Performance: Webhooks delegate work to separate worker nodes, resulting in better performance and preventing the main N8N instance from becoming unresponsive.
  • Response Data: Webhooks can return response data (execution ID), which is essential for tracking executions and updating the Q table.

However, webhooks may count towards the monthly execution allowance in cloud-based N8N plans, while subworkflows may not.

Data Passed to Webhook

The following data is passed to the rag ingestion workflow via the webhook:

  • ID: The ID of the file.
  • Name: The name of the file.
  • Path: The path to the file on the SFTP server.
  • First Attempt: A boolean indicating whether this is the first attempt to process the file.

Rag Ingestion Workflow Details

The rag ingestion workflow:

  1. Downloads the file: From the SFTP server.
  2. Processes the file: Chunks the data, creates embeddings, and stores them in the vector database.
  3. Updates the Q table: To indicate the status of the execution (success).
  4. Moves the file to a processing folder: Only on the first attempt.

Key Lessons Learned

The video concludes with five key lessons learned during the process of scaling N8N workflows:

  1. Turn off saving execution data for successful executions: To prevent database bloat, especially when handling large amounts of data.
    • Set the executions_data_save_on_success environment variable to none.
    • Only turn this off when scaling up, as it makes debugging more difficult.
  2. Use multiple layers of error handling:
    • Enable "Retry on Fail" for external API services.
    • Implement a catch-all error handler workflow to automatically retry executions.
    • Ensure the workflow can gracefully handle upserts to prevent duplication.
  3. Don't overload executions with too much data at once:
    • Break up workflows into smaller, separate workflows.
    • Use webhooks to call separate workflows and delegate work to worker nodes.
  4. Optimize Superbase settings:
    • Upgrade to a Pro account for increased database space.
    • Upgrade to at least a Micro compute instance.
    • Monitor resource usage and increase compute size if necessary.
  5. Building the orchestrator is only one piece of the puzzle:
    • Ensure the N8N instance can handle the scale (multiple workers, Q mode).
    • Tune the workflows for performance and efficient resource utilization.
    • Benchmark performance and track changes to identify bottlenecks.

Conclusion

The video provides a detailed guide to building an orchestrator workflow for scaling rag ingestion in N8N. By parallelizing execution, handling errors gracefully, and optimizing both the orchestrator and ingestion workflows, it's possible to reliably process thousands of files and build a robust RAG system. The key takeaways are the importance of parallel processing, error handling, data management, and continuous optimization.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.