Key Concepts
- Local AI Agents
- Self-Hosting
- Offline Functionality
- Notebook LM (Insights LM)
- Superbase
- N8N
- Ollama
- Whisper ASR (Automatic Speech Recognition)
- KIT TTS (Text-to-Speech)
- Vector Store
- LLM Inference
- Docker Containers
- Edge Functions
- Model Quantization
- GPU VRAM
Local Insights LM Demo and Setup Guide
Introduction
The video demonstrates a fully local version of Insights LM, a self-hosted application similar to Notebook LM, that allows users to chat with their documents. This version is designed to run completely offline, addressing privacy concerns, reducing cloud dependency, and cutting costs. The system leverages Superbase, N8N, Ollama, Whisper ASR, and KIT TTS, all running within Docker containers.
Architectural Changes
To achieve local functionality, significant architectural changes were made to the backend flows, primarily within N8N. The front-end app required no code changes due to the flexibility of N8N as a backend. The local AI package created by Cole Medan served as the foundation for this project.
System Components and Functionality
- Docker Desktop: Hosts multiple containers, including Insights LM, Superbase, N8N, Ollama, Whisper ASR, and KIT TTS.
- Superbase: Provides the backend database and vector store for document embeddings.
- N8N: Acts as the workflow automation engine, orchestrating the various processes.
- Ollama: Handles LLM inference and embeddings, utilizing the Quinn 3B 8 billion parameter model with quantization 4.
- Whisper ASR: Transcribes audio files (e.g., MP3) into text.
- KIT TTS: Generates speech from text for the deep dive conversation feature.
Notebook Creation and Embedding Process
- Upload Sources: Users can upload various file types (PDF, MP3, text) or provide URLs.
- Generate Notebook Details Workflow: N8N workflow triggered to generate the notebook title and description using Ollama.
- Upsert to Vector Store: The document is chunked and embedded into Superbase's vector store using the nomic embed text model.
- Performance: Embedding a 180-page document took approximately 8 seconds using the nomic embed text model.
Chat Workflow
- Question Input: User asks a question related to the uploaded document.
- Chat Workflow Trigger: N8N workflow activated.
- Search Query Generation: Quinn 3 model generates a search query based on chat history.
- Vector Database Retrieval: Relevant chunks are retrieved from the vector store.
- Response Generation: Ollama uses the retrieved chunks to generate a response.
- Citation Processing: Citations are extracted and formatted into a JSON structure for clickable links to the document sections.
- Local LLM Adaptation: The chat flow was re-engineered for local LLMs due to the limited reasoning capabilities of smaller models compared to frontier models. The process was broken down to reliably produce responses and citations.
Additional Features
- Audio Transcription: Uploading an MP3 file triggers the Whisper transcription container, providing a summary and full transcription.
- URL Scraping: The system can scrape content from URLs and generate summaries.
- Note Saving: Users can save notes within the notebook.
- Deep Dive Conversation: Generates a podcast-style conversation using KIT TTS, though the quality is noted to be less sophisticated than cloud-based solutions like Gemini's text-to-speech.
Setup Guide: Step-by-Step
- Prerequisites:
- Python
- Git or GitHub Desktop
- Docker or Docker Desktop
- VS Code (recommended)
- Clone Repositories:
- Clone Cole Medan's
local-airepository. - Clone the Insights LM local package repository into the same folder.
- Clone Cole Medan's
- Configure Environment Variables:
- Copy
env.exampleto.env. - Generate and set secret keys for N8N, Postgres, JWT, anonymous, and service roles.
- Set the
NOTEBOOK_GENERATION_ALT_KEYfor webhooks. - Copy the content of
insights-lm-local-package/superbase/.envto the end of.env.
- Copy
- Modify Docker Compose File:
- Copy the
whisper_cachevolume definition frominsights-lm-local-package/docker-compose-copy.ymlto the maindocker-compose.yml. - Copy the
insights-lm,kitti_texttospeech, andwhisper_asrservice definitions frominsights-lm-local-package/docker-compose-copy.ymlto theservicessection of the maindocker-compose.yml. - Update the Quinn model to "TheBloke_OpenHermes-2.5-Mistral-7B-GGUF/openhermes-2.5-mistral-7b.Q4_K_M.gguf" in the docker-compose file.
- Copy the
- Start Services:
- Use the
start-services.pyscript with the appropriate profile (e.g., Nvidia GPU).
- Use the
- Superbase Setup:
- Access the Superbase Studio dashboard via the exposed port.
- Run the Superbase migration script from
insights-lm-local-package/superbase/migrations/schema.sqlin the SQL editor.
- Superbase Functions:
- Move the function folders from
insights-lm-local-package/superbase/functionstosuperbase/volumes/functions. - Modify the
superbase/docker/docker-compose.ymlfile to include the environment variables frominsights-lm-local-package/superbase/docker-compose.ymlin thefunctionsservice definition.
- Move the function folders from
- Restart Services:
- Stop and restart the services to apply the changes.
- N8N Setup:
- Access N8N via the exposed port.
- Create an account.
- Import the workflows from
insights-lm-local-package/n8n/import-workflows.json. - Update the N8N API key in the "n8n API Request" node.
- Enter the Superbase credential ID, webhook header authentication credential ID, and Lama credential ID in the "Enter User Values" node.
- Activate all workflows except the "Extract Text" workflow.
- Insights LM Access:
- Access Insights LM via the exposed port (3000).
- Create a user in Superbase authentication.
- Log in to Insights LM and start creating notebooks.
Hardware Considerations
Running local models requires sufficient hardware resources. The video demonstrates the system running on an Nvidia GeForce 4070 with 8GB of VRAM. However, larger models may require more VRAM. The choice of model should be appropriate for the available hardware.
Conclusion
The video provides a comprehensive guide to setting up a fully local, offline version of Insights LM. While it may not match the sophistication of cloud-based solutions, it offers greater control over data privacy and reduces reliance on external infrastructure. The step-by-step instructions and detailed explanations make it possible for users to deploy this system on their own machines.
AI summaries can miss context or contain errors. Check important details against the original video.