LocalGPT 2.0: Turbo-Charging Private RAG

Prompt EngineeringAbout 4 min readJun 20, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Local GPT: A private, open-source project for chatting with documents without external APIs.
  • RAG (Retrieval-Augmented Generation): A framework for enhancing LLMs with external knowledge.
  • Triage Agent/Classifier: A component that determines whether to use RAG, internal knowledge, or chat history to answer a query.
  • Contextual Retrieval: A technique to add context from surrounding chunks to a given chunk.
  • Structure-Aware Chunking: Chunking that preserves the structure of documents (e.g., titles, sections).
  • Multi-Vector Representation: Using multiple embedding models for a more robust representation of data.
  • Reasoning Model: An LLM used to generate answers for subqueries.
  • Verifier: An independent component that grades the final answer for accuracy.
  • LanceDB: A single-file database used for storing embedding vectors.

1. Introduction to the New Local GPT

  • The video introduces a new version of Local GPT, an open-source project that allows users to chat with their documents privately.
  • The new version is written from scratch in pure Python, without relying on frameworks like LangChain or LlamaIndex.
  • It boasts features not commonly found in open-source RAG implementations.
  • The presenter will demonstrate the system in action and then discuss the technical details.

2. Demonstration of Local GPT in Action

  • The presenter selects an index containing invoices and the DeepSeek paper.
  • A query is posed: "What is the invoice amount that this company sent to Horizon Media and where is Solar Wind Energy located?"
  • Local GPT analyzes the query, decides whether to use the RAG pipeline or its internal knowledge, and creates subqueries if necessary.
  • It performs retrieval, reranking, and context window expansion.
  • A reasoning model generates answers for each subquery in parallel.
  • The answers are combined with the user input to generate a final response.
  • A verifier assigns a score to the final response.

3. Architecture Overview

  • The original Local GPT had a naive RAG pipeline.
  • The new version has a more complex architecture, which the presenter will explain in detail.
  • The implementation is divided into three parts:
    • Front End: Still in progress, with some bugs.
    • Back End: An agentic workflow that decides when to use which step.
    • Store Components: Stores embedding vectors, keyword search index, chat history, and caching.
  • The system is currently hosted through Ollama, but VLLM integration is planned for production systems.

4. Indexing and Vector Store Creation

  • Currently supports PDF files, with plans to support other unstructured data types and structured data.
  • PDF files are converted to Markdown using Docling to preserve document structure.
  • Structure-Aware Chunking: Chunks are created considering the structure of the Markdown file (chunk size is about 500 tokens).
  • Two key processes are performed:
    • Document-Level Summary: The first five chunks of each document are used to create an overview summary, stored in a database.
    • Contextual Retrieval: Each chunk is contextualized using a sliding window of surrounding chunks, and a summary is appended to the chunk.
  • The contextualized chunks are stored in a vector store (LanceDB).
  • A keyword index is also extracted for each chunk.
  • The presenter uses the Quen 3 embedding model, but plans to implement a more robust multi-vector representation.

5. Retrieval Process

  • When a user query comes in, it first hits the triage agent/classifier.
  • The triage agent makes three decisions:
    • Whether to use the RAG pipeline.
    • Whether to answer using internal knowledge.
    • Whether to answer using chat history.
  • The triage agent uses a small LLM (6 billion model) and has access to the document overview summaries.
  • The system decides whether to decompose the query into subqueries.
  • For each subquery, it performs retrieval on the dense embedding model and BM25.
  • Top chunks are passed through a cross-encoder for reranking (ColBERT-style reranker).
  • Context Window Expansion: The window around the reranked chunks is expanded (two chunks around it).
  • A reasoning model analyzes each chunk individually and generates answers for each subquery.
  • All context, answers, and the original question are passed to a final LLM to generate the final answer.
  • An independent verifier grades the final answer based on the original context and queries.
  • The system assigns a confidence level to the answer.

6. Key Arguments and Perspectives

  • There is no one-size-fits-all RAG solution; every solution is application-dependent.
  • Chunking and data processing are highly dependent on the type of information being fed into the system.
  • The presenter is looking for design partners to create niche-specific versions of Local GPT.

7. Development Process

  • The code for Local GPT was written using a combination of AI tools, including Gemini, OpenAI, 3 Cloud Cursor, and Cloud Code.
  • The presenter prefers to implement things in small pieces and gradually build on top of them.

8. Conclusion

  • The new Local GPT is a complex and robust RAG system with several advanced features.
  • The presenter plans to release the new version in a few weeks.
  • The original Local GPT remains available on GitHub.
  • The presenter encourages contributions to the open-source project.
  • The presenter offers consulting and advising services for businesses looking for scalable RAG solutions.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.