Build with n8n + Ollama (DeepSeek R1, CodeLlama) to Route Prompts and Cut AI Spend

By F5 DevCentral Community

Share:

Key Concepts

  • Mixture of Experts (MoE) Architecture: An AI design where a "router" or "gate" model directs incoming tasks to one or more specialized "expert" models, optimizing resource use and performance.
  • AI Agents: Software components designed to perform specific tasks, often leveraging AI models, in this context, specialized for reasoning, coding, or general queries.
  • N8N: A powerful workflow automation tool used to visually construct and manage complex processes, including those involving AI agents and models.
  • Olama: A framework for running large language models (LLMs) locally, providing an API for integration into applications.
  • Docker: A containerization platform used to package applications and their dependencies into isolated environments, ensuring consistent deployment.
  • Text Classifier: An AI agent responsible for analyzing incoming text and assigning it to predefined categories (e.g., reasoning, coding, general).
  • Reasoning Agent: An AI agent specialized in handling prompts that require deep understanding, analysis, or complex problem-solving.
  • Coding Agent: An AI agent specialized in generating or assisting with code, particularly for specific languages or frameworks (e.g., Python, JSON, Node.js, I rules).
  • General Agent: A fallback AI agent designed to handle prompts that do not fit into more specialized categories, providing broad knowledge.
  • LLMs (Large Language Models): AI models like DeepSeek R1, Llama 3.2, and Code Llama, used as the underlying intelligence for the agents.
  • Tokens: The fundamental units of text processed by LLMs, directly impacting computational cost and speed.
  • System Prompts: Specific instructions given to an LLM to define its role, constraints, and desired output format, crucial for guiding agent behavior.

Comprehensive Summary of AI Step-by-Step Lab: Multi-Model AI Agents with N8N and Olama

This lab demonstrates the construction of a sophisticated AI application using N8N to orchestrate multiple AI agents, leveraging a Mixture of Experts (MoE) architecture. The primary goal is to enhance the efficiency of AI prompt processing and reduce operational costs by intelligently routing prompts to specialized, cost-effective models, reserving larger, more resource-intensive reasoning models only when absolutely necessary.

1. Core Architecture and Efficiency Rationale

The system employs an MoE architecture where a central Text Classifier agent acts as a router. This classifier analyzes incoming prompts and directs them to the most appropriate specialized AI agent. This approach ensures that a "real larger scale reasoning model" is only engaged when the prompt genuinely requires deep intelligence, thereby minimizing token usage and infrastructure costs. The architecture defines three main types of specialized agents: a Reasoning Agent, a Coding Agent, and a more General Agent.

2. Prerequisites and Environment Setup

Participants are expected to have completed prior Dev Central labs on Olama installation and N8N setup (either on Mac or Docker). The lab assumes a two-machine setup: one dedicated to running Olama with downloaded LLM models (the "model server") and another for running N8N (the "application server"). Both servers require Docker installed.

3. Step-by-Step Implementation Process

The lab outlines a detailed process for setting up the environment and building the N8N workflow:

  • LLM Host Configuration:
    • Configure the LLM host to utilize NVIDIA hardware.
    • Restart the Docker daemon to apply changes.
  • Olama Setup:
    • Create a Docker volume for persistent model data: docker volume create olama-model-data.
    • Run the Olama Docker container, mapping the created volume and exposing port 11434: docker run -d -v olama-model-data:/root/.ollama --name olama -p 11434:11434 ollama/ollama.
    • Install specific LLM models using docker exec olama ollama pull [model_name]:
      • DeepSeek R1 1.5 billion: Used for the text classifier and general agent.
      • DeepSeek R1 7 billion: The "real smarts" model, used for the reasoning agent.
      • Llama 3.2: An alternative general agent model.
      • Code Llama: A highly refined model explicitly for coding tasks (Python, JSON, Node.js, I rules).
  • N8N Setup:
    • Create an N8N data volume for persistent workflows: docker volume create n8n-data.
    • Pull and run the latest N8N Docker container, mapping the volume and exposing port 5678: docker run -d -v n8n-data:/home/node/.n8n --name n8n -p 5678:5678 n8n/n8n.
    • Access the N8N web UI via https://[app_server_IP]:5678.
    • Initial N8N Configuration: Set up an owner account with a valid email (for license key delivery), skip customization, and activate the community edition using the emailed license key in the "Settings" section.

4. Building the N8N Workflow (Canvas)

The core of the application is built visually within N8N's canvas:

  • Trigger Node: A "Chat Message Received" node is added as the entry point for incoming prompts, simulating a chatbot in a Slack channel.
  • Test Data Generation: A test chat message (e.g., "Which language is better, C++ or Java?") is sent to generate input data for workflow development.
  • Text Classifier Node:
    • Configured to classify the JSON.input from the chat message.
    • Categories Defined:
      • Reasoning: "if reasoning need is indicated by the chat message, this is the category to assign."
      • Coding: "if the chat message indicates a need to code or a asks for help with computer languages and scripting language scripting languages like I rules, JSON or NodeJ.js. Assign this category."
    • Options: "Output on an extra other branch" when no clear match, ensuring prompts are not discarded.
    • System Prompt: A critical system prompt is used to guide the classifier: "Please classify the text provided by the user into one of the following categories: [category string]. If they explicitly ask for coding help, do not fail and classify the message as coding. If they explicitly ask for reasoning help, do not fail and classify the message as reasoning. Otherwise, send the JSON chat input." This prompt also specifies "do not allow multiple cases to be true."
    • Olama Model Connection: An "Olama Model" node is connected to the classifier. New credentials are created pointing to the Olama server (10.1.1.5:11434). The DeepSeek R1 1.5 billion model is selected for its reasoning capabilities in classification.
    • Retry Mechanism: The Text Classifier node is configured to "Retry on fail" with 3 attempts, addressing potential connection or cloud-related issues.
  • Specialized AI Agents:
    • Reasoning Agent: An "AI Agent" node, renamed "Reasoning Agent," is configured with retry on fail (3 tries). It connects to the Olama server and utilizes the DeepSeek R1 7 billion model for its advanced reasoning capabilities.
    • Coding Agent: Another "AI Agent" node, renamed "Coding Agent," is configured with retry on fail (3 tries). It connects to Olama and uses the Code Llama (latest) model, specifically designed for code generation.
    • General Agent (Fallback): A third "AI Agent" node (implicitly for general queries) is configured with retry on fail (3 tries). It connects to Olama and uses either DeepSeek R1 1.5 billion or Llama 3.2 as a decent general-purpose model.

5. Workflow Testing and Validation

The lab demonstrates testing the complete workflow by sending a specific coding-related prompt: "I need help coding an application written in Python. Could you please provide a snippet of code that will post pictures to LinkedIn, please?" The execution flow is observed: the chat message triggers the Text Classifier (using DeepSeek R1 1.5 billion), which correctly identifies the prompt as "coding" and routes it to the Coding Agent. The Coding Agent then uses Code Llama to generate and return a code snippet.

6. Key Arguments and Future Considerations

  • Non-Deterministic Nature of AI: The presenter acknowledges that AI is "non-deterministic technology," meaning outputs can vary. The goal is to "make it where it doesn't vary that much" through careful tuning.
  • Tuning and Reliability: Emphasizes the importance of system prompts and retry mechanisms to improve the consistency and reliability of AI responses, especially when dealing with classification failures or ambiguous inputs.
  • Model Quality: While the demonstrated free models provide a sound architectural proof-of-concept, the presenter notes that using higher-end commercial models (e.g., Claude, ChatGPT) or more powerful local models (with sufficient hardware) would yield better quality and more consistent results.
  • Open-Ended Development: The lab encourages further experimentation, such as integrating the chatbot with platforms like Slack, and exploring different types of prompts (e.g., political questions, simple general knowledge) to observe how the agents classify and respond.

7. Synthesis and Conclusion

This lab provides a practical, detailed guide to building an intelligent, cost-efficient AI application using N8N and Olama. By implementing a Mixture of Experts architecture, the system effectively routes diverse prompts to specialized AI agents powered by appropriate LLMs, optimizing resource utilization and reducing operational costs. The emphasis on system prompts, retry mechanisms, and iterative testing highlights best practices for developing robust AI solutions, even with the inherent non-deterministic nature of the technology. The framework offers a flexible foundation for creating advanced chatbots capable of handling a wide range of queries with improved efficiency and targeted intelligence.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video