Key Concepts
- RAG (Retrieval-Augmented Generation): A framework for enhancing LLMs by grounding them in external data sources.
- Vector Store: A database that stores data as vector embeddings for efficient similarity search.
- Embeddings: Numerical representations of text or other data, capturing semantic meaning.
- Chunking: Dividing documents into smaller segments for indexing and retrieval.
- Metadata Filtering: Using metadata associated with documents to refine search results.
- Hybrid Search: Combining keyword-based search with semantic search for improved retrieval.
- Re-ranking: Ordering retrieved chunks by relevance using a cross-encoder model.
- Contextual Retrieval: Adding context to chunks to improve retrieval accuracy.
- Prompt Caching: Storing and reusing LLM responses to reduce costs and latency.
- Late Chunking: Embedding the entire document before chunking to preserve context.
- Edge Functions: Serverless functions deployed close to users for low-latency execution.
- Data Ingestion Pipeline: Automated workflow for loading and processing data into a vector store.
- Record Manager: System for tracking documents in a vector store to prevent duplication.
- Web Scraping: Extracting data from websites for ingestion into a vector store.
- OCR (Optical Character Recognition): Converting scanned images of text into machine-readable text.
RAG Master Class Overview
The AI Automators NAND RAG Master Class is a practical, hands-on course focused on building RAG systems. It emphasizes building solutions quickly while learning key concepts. The course covers beginner, intermediate, and advanced topics, with the beginner and intermediate sections being no-code. Some vibe coding with an AI model is required for the advanced section.
Course Structure
The course is structured across seven lessons:
- Lesson 1: Introduction to RAG using Notebook LM and OpenAI's Assistants API.
- Lesson 2: Building a simple RAG agent using NADN's AI Agent feature and Superbase vector store.
- Lesson 3: Setting up a RAG pipeline to ingest documents into a vector store.
- Lesson 4: Extending the pipeline to handle different file types and formats, including OCR for scanned PDFs.
- Lesson 5: Integrating web scraping using Firecrawl AI into the RAG pipeline.
- Lesson 6: Implementing hybrid search to combine keyword and semantic search. Implementing re-ranking with coheres 3.5 reranker model.
- Lesson 7: Implementing contextual retrieval to improve retrieval accuracy.
Lesson 1: Introduction to RAG
- Main Topics: Demonstrating what RAG is and setting up a simple RAG agent.
- Key Points:
- RAG grounds LLMs in external data sources.
- Notebook LM is a web app with a built-in RAG engine.
- OpenAI's Assistants API allows programmatic interaction with RAG systems.
- It's crucial to instruct the LLM to avoid speculating and to say "Sorry, I don't know" if the answer isn't in the provided information.
- Examples:
- Using Formula 1 regulations as a data source.
- Comparing the responses of Notebook LM and OpenAI's Assistants API.
- Process:
- Upload documents to Notebook LM or OpenAI's Assistants API.
- Ask questions related to the documents.
- Observe how the system retrieves and generates answers based on the provided data.
- Limitations of OpenAI's Assistants API: Lack of transparency in chunk retrieval and citations.
Lesson 2: Building a Simple RAG Agent with NADN
- Main Topics: Using NADN's AI Agent feature and Superbase vector store to chat with documents.
- Key Points:
- NADN's AI Agent feature allows creating agents with memory, knowledge, and tools.
- Superbase can be used as a persistent vector store.
- Prompt engineering is crucial to convince the LLM not to respond if it's not getting any relevant information back from the Rag store.
- The simple vector store is not for production use because it saves the data in memory.
- Process:
- Create an AI Agent in NADN.
- Configure the agent with a system message and chat model.
- Add a vector store tool to the agent.
- Load documents into the vector store.
- Test the agent by asking questions related to the documents.
- Technical Terms:
- Temperature: A parameter that controls the randomness of the LLM's responses.
- Context Window Length: The number of recent exchanges that the agent remembers.
- Embedding Model: A model that converts text into vector embeddings.
- Text Splitter: A tool that divides documents into smaller chunks.
- Superbase Setup:
- Create a Superbase account and project.
- Set up a vector database using the Langchain quick start script.
- Create an API key for the Superbase project.
- Create a connection in NADN using the Superbase host and service role secret.
Lesson 3: Building a Data Ingestion Pipeline
- Main Topics: Building a data ingestion pipeline to load multiple documents into a vector store and implementing a record manager to avoid duplicate documents.
- Key Points:
- Using Google Drive as a source for documents.
- Implementing a record manager to track documents in the vector store.
- Using a crypto node to generate a unique hash of each file.
- Using a switch node to handle different scenarios based on whether the document exists in the record manager and whether the hashes are the same.
- Process:
- Set up a Google Drive trigger to watch for changes in a specific folder.
- Loop through each file that has been provided by the trigger.
- Download the file.
- Generate a unique hash of the file.
- Search the record manager to see if the file already exists.
- If the file doesn't exist, create a row in the record manager and upsert the document to the vector store.
- If the file exists and the hashes are the same, move on to the next file.
- If the file exists and the hashes are different, delete the old vectors, update the record manager, and upsert the new vectors.
- Technical Terms:
- SHA 256: A cryptographic hash function.
- Upsert: Insert or update a row in a database.
- Logical Connections: The data ingestion pipeline is connected to the AI agent through the vector store.
Lesson 4: Handling Different File Types and Formats
- Main Topics: Expanding the RAG pipeline to handle different file types and formats, including OCR for scanned PDFs.
- Key Points:
- Using a switch node to handle different file formats.
- Using the extract from file node to extract text from different file formats.
- Using the Google Docs get document node to extract text from Google Docs.
- Using the HTML to markdown node to convert HTML to markdown.
- Using Mistral OCR to extract text from scanned PDFs.
- Process:
- Set up a switch node to handle different file formats.
- For each file format, use the appropriate node to extract the text.
- Generate a unique hash of the text.
- Search the record manager to see if the file already exists.
- If the file doesn't exist, create a row in the record manager and upsert the document to the vector store.
- If the file exists and the hashes are the same, move on to the next file.
- If the file exists and the hashes are different, delete the old vectors, update the record manager, and upsert the new vectors.
- Mistral OCR Integration:
- Create a Mistral account and API key.
- Use the HTTP request node to upload the PDF file to Mistral OCR.
- Use the HTTP request node to get a signed URL for the uploaded file.
- Use the HTTP request node to trigger the OCR process.
- Aggregate the markdown fields from each page into a single text string.
Lesson 5: Integrating Web Scraping
- Main Topics: Integrating web scraping using Firecrawl AI into the RAG pipeline.
- Key Points:
- Using Firecrawl AI to crawl and scrape websites.
- Adapting the existing ingestion workflow to work for both web scraping and document uploads.
- Using a web hook to receive data from Firecrawl.
- Abstracting the file ID to a document ID to handle both files and web pages.
- Process:
- Set up a Firecrawl account and API key.
- Use the HTTP request node to trigger a crawl on Firecrawl.
- Set up a web hook to receive data from Firecrawl.
- Loop through each page that has been crawled.
- Extract the text from the page.
- Generate a unique hash of the text.
- Search the record manager to see if the page already exists.
- If the page doesn't exist, create a row in the record manager and upsert the page to the vector store.
- If the page exists and the hashes are the same, move on to the next page.
- If the page exists and the hashes are different, delete the old vectors, update the record manager, and upsert the new vectors.
- Technical Terms:
- Web Hook: A mechanism for one application to notify another application when an event occurs.
Lesson 6: Implementing Hybrid Search and Re-ranking
- Main Topics: Implementing hybrid search to combine keyword and semantic search and re-ranking to order retrieved chunks by relevance.
- Key Points:
- Hybrid search combines full text search with semantic search for improved retrieval.
- Re-ranking orders retrieved chunks by relevance using a cross-encoder model.
- Coher 3.5 is a popular re-ranking model.
- Superbase edge functions and database functions are used to implement hybrid search.
- Process:
- Set up hybrid search in Superbase using edge functions and database functions.
- Generate embeddings for the query using OpenAI's embeddings API.
- Call the Superbase edge function to perform hybrid search.
- Send the retrieved chunks to the Coher 3.5 re-ranker API.
- Reorder the chunks based on the re-ranker's output.
- Return the reordered chunks to the AI agent.
- Technical Terms:
- Reciprocal Ranked Fusion: An algorithm for combining the results of keyword search and semantic search.
- Cross-Encoder Model: A neural network that compares two texts and outputs a relevance score.
Lesson 7: Implementing Contextual Retrieval
- Main Topics: Implementing contextual retrieval to improve retrieval accuracy.
- Key Points:
- Contextual retrieval adds context to chunks to improve retrieval accuracy.
- Anthropic popularized contextual retrieval in their contextual retrieval paper.
- OpenAI's prompt caching is used to reduce costs and latency.
- Process:
- Chunk the document manually using a code node.
- Loop through each chunk.
- Use OpenAI's message and model API to generate a description for each chunk.
- Combine the chunk description with the chunk text.
- Upsert the combined text to the vector store.
- Technical Terms:
- Prompt Caching: Storing and reusing LLM responses to reduce costs and latency.
Conclusion
The RAG Master Class provides a comprehensive guide to building RAG systems using NADN. It covers various techniques for improving the accuracy and quality of retrieval, including hybrid search, re-ranking, and contextual retrieval. The course emphasizes practical implementation and provides step-by-step instructions for building a complete RAG pipeline.
AI summaries can miss context or contain errors. Check important details against the original video.