To Scale our RAG Agent (5,000 Files per/hr)

The AI AutomatorsAbout 7 min readSep 22, 2025Watch original
THE SUMMARYAI-generated

N8N Rag Workflow Tuning and Scaling: A Deep Dive

Key Concepts:

  • Rag (Retrieval-Augmented Generation): An AI framework where a model retrieves information from an external knowledge base before generating a response.
  • N8N: A no-code workflow automation platform.
  • Superbase: An open-source Firebase alternative, used here as a vector store.
  • Vector Store: A database that stores data as vector embeddings for similarity search.
  • Orchestrator: A custom queue system built to manage and parallelize file processing.
  • Bottleneck: A point in a workflow that limits overall performance.
  • Vertical Scaling: Increasing resources (CPU, RAM) on a single server.
  • Horizontal Scaling: Adding more servers to distribute the workload.
  • Workers: N8N instances dedicated to processing tasks from a queue.
  • Queue Mode: N8N configuration where tasks are added to a queue (Redis) for workers to process.
  • Main Mode: N8N configuration where the main instance handles all tasks.
  • Concurrency: The number of tasks a worker can process simultaneously.
  • Mistral OCR: An Optical Character Recognition (OCR) service used to extract text from images in PDF files.
  • Contextual Embeddings: Enriching data chunks with snippets providing context within the document.

1. Introduction and Goal

The video details a 100-hour effort to optimize and scale an N8N Rag workflow designed to import files into a Superbase vector store for AI agent querying. The initial workflow was slow, processing only around 100 files per hour. The goal was to dramatically increase throughput, aiming for 5,000 PDF files per hour (10GB of data, 250,000 chunks).

2. Approach: Tuning and Orchestration

The project was divided into two main tasks:

  • Tuning (Speaker's Task): Optimizing the existing Rag ingestion pipeline to identify and eliminate bottlenecks.
  • Orchestration (Allan's Task): Building a custom queue system (orchestrator) to parallelize file processing and ensure resilience with automatic retries. Allan's video on the orchestrator is linked in the description.

3. Tuning Methodology: Iterative Approach

The tuning process followed an iterative approach:

  1. Establish a Benchmark: Measure the baseline performance of the workflow.
  2. Identify Bottlenecks: Pinpoint the slowest nodes in the workflow.
  3. Generate a Hypothesis: Formulate potential solutions to address the bottlenecks.
  4. Experiment: Implement the proposed changes.
  5. Test: Evaluate the impact of the changes on performance.
  6. Analyze Results: Determine whether the changes improved performance.
  7. Repeat: Iterate through the process, refining the workflow.

4. Establishing a Baseline

A single PDF file containing 50,000 characters (a 24-page coffee machine instruction booklet) was used as the benchmark. The initial processing time, including Mistral OCR, was 26 seconds per file. This translated to approximately 120 files per hour.

5. Identifying Bottlenecks: Execution History and Logs

N8N's execution history and logs were used to identify bottlenecks. By examining the time taken by each node, the speaker identified several areas for improvement:

  • Downloading files from Google Drive (2 seconds)
  • Uploading to Mistral (1.5 seconds)
  • Extracting metadata with OpenAI's GPT4 mini (1.7 seconds)
  • Upserting data into Superbase (multiple instances, 3-4 seconds each)

6. Tuning Experiments and Results

6.1. Increasing Batch Size in the Main Loop

  • Hypothesis: Processing multiple files in the main loop (instead of one at a time) would improve throughput.
  • Implementation: Modified the workflow to process files in batches of 5, 10, and 20. This involved using merge nodes to manage data flow due to nested loops and mappings.
  • Results: A small improvement (10%) was observed with a batch size of 5. Larger batch sizes (10, 20) either didn't improve performance or crashed the server due to excessive RAM usage when downloading multiple binary files.
  • Decision: Initially discarded, but later re-evaluated.

6.2. Increasing Chunk Batch Size

  • Hypothesis: Increasing the batch size for chunk processing would reduce the number of calls to OpenAI for contextual embedding generation.
  • Implementation: Increased the chunk batch size from 20 to 200.
  • Results: A 20% improvement in processing speed was achieved.
  • Decision: The change was kept.

6.3. Replacing the Vector Store Node

  • Hypothesis: The Superbase vector store node was a major bottleneck.
  • Investigation: The logs showed that upserting chunks into Superbase was taking a significant amount of time. Tests revealed that the custom chunking process, when used with the vector store node, was the primary cause of the slowness.
  • Implementation: Replaced the vector store node with a direct SQL query into the Superbase documents table. This query was created with the help of ChatGPT.
  • Results: A 55% improvement in processing speed was achieved. The time to insert data into the database was reduced to less than a second.
  • Decision: The change was kept.

7. Scaling the Infrastructure: Vertical and Horizontal Scaling

  • Vertical Scaling: Adding resources (CPU, RAM) to a single server. N8N, being a NodeJS application, is single-threaded and limited to one CPU core.
  • Horizontal Scaling: Adding more servers to distribute the workload. This is achieved using workers and queuing.
  • Workers and Queuing: N8N instances are configured as workers, which pull tasks from a Redis queue and process them. The main N8N instance adds tasks to the queue.
  • Recommendation: Use Q mode for scaling N8N.

8. Scaling with Railway

Railway is recommended as a platform for easily testing N8N scaling without manual server configuration. Railway provides a template with N8N main, a worker, a Postgres database, and a Redis queue. Workers can be easily duplicated and redeployed. Railway also uses replicas, which are multiple instances of a worker behind a load balancer.

9. The Orchestrator: Managing Parallel Executions

The orchestrator is a custom N8N workflow designed to:

  • Fetch files from a server.
  • Trigger the Rag ingestion flow for each file in parallel using webhooks.
  • Manage a queue on top of the Redis queue to track execution status (success/failure).
  • Implement automatic retries for failed executions.

10. Performance Results with Orchestration and Scaling

  • With two workers, a 55% improvement in processing speed was achieved (4.1 seconds per file).
  • The speaker initially expected a larger improvement with more workers due to the concurrency parameter (default 10), which allows each worker to process multiple tasks simultaneously.
  • The bottleneck was identified in the orchestrator: files were being moved to a processing folder sequentially before import.
  • Disabling the sequential file moving in the orchestrator reduced the processing time to 1.1 seconds per file.
  • Increasing the batch size to 50 files (with four workers and replicas) further reduced the processing time to 0.7 seconds per file.
  • Re-enabling Mistral OCR increased the processing time to 1.4 seconds per file, which was still considered excellent.

11. Lessons Learned

  1. Systematic Approach: Use a structured approach (e.g., a spreadsheet) to track configuration settings and results. Dive deep into the logs to identify bottlenecks.
  2. Custom Job: General best practices are helpful, but you need to immerse yourself in the specific workflow, data, and third-party integrations.
  3. Resources Aren't Always the Solution: Increasing resources (e.g., workers) can sometimes slow things down. Analyze the flow to understand the underlying issues.
  4. Third-Party Rate Limiting: Third-party APIs can rate limit or throttle requests.
  5. Go Direct to the Source: Consider bypassing APIs and connecting directly to the data source (e.g., using a direct SQL query to Superbase).
  6. Focus on Low-Hanging Fruit: Start with easier optimizations before tackling complex changes.
  7. Know When to Quit: There is a law of diminishing returns. Aim for a balance between performance and resilience.
  8. Self-Hosting: As you scale, consider self-hosting N8N, Superbase, and other services.

12. Conclusion

The video demonstrates a successful effort to significantly improve the performance of an N8N Rag workflow. By systematically identifying and addressing bottlenecks, implementing horizontal scaling, and building a custom orchestrator, the team achieved a 97% reduction in processing time, enabling the import of 5,000 PDF files per hour. The lessons learned provide valuable insights for anyone looking to optimize and scale their own N8N workflows.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.