Key Concepts
AI Coding Assistants, Hallucination, RAG (Retrieval Augmented Generation), MCP (Multi-Context Pipeline) Server, Archon, Boomerang Tasks, Rue Code, Context 7, Cloud Task Manager, Knowledge Backbone, Contextual Embeddings, Hybrid Search, Agentic RAG, Re-ranking, Cross-Encoder Model, Vector Database, Semantic Search, Keyword Search, LLM (Large Language Model), Superbase, Versell AI SDK, OpenAI, N8N, Dynamis AI Mastery.
Crawl for AI RAG MCP Server: Enhancements and Future Vision
Introduction
The video discusses the challenges of AI coding assistants hallucinating and proposes a solution: the Crawl for AI RAG MCP server. The goal is to provide AI coders with up-to-date documentation and a robust knowledge backbone, ultimately evolving into the next iteration of Archon.
Problem Statement
AI coding assistants often struggle with accuracy due to a lack of comprehensive understanding of project context and documentation, leading to hallucinations.
Proposed Solution: Crawl for AI RAG MCP Server
The MCP server aims to address this by:
- Providing a strong knowledge backbone using RAG.
- Enabling different agents to manage various project aspects.
- Facilitating task planning and management.
Setup and Configuration
The MCP server is open-source and free to use. The video provides a link to the server and instructions for configuration, including enabling/disabling different RAG strategies. The setup involves connecting to the server within an AI coding assistant (e.g., Windsurf) and configuring Superbase for knowledge storage.
Data Ingestion and Knowledge Base
The server supports crawling various website types:
- LLM's text (single-page documentation).
- Recursive scraping of websites via base page navigation.
- Sitemaps.
Example: The video uses the Versell AI SDK documentation (specifically the LLM.ext page) as a demo.
The crawling process involves:
- Copying the documentation link.
- Using the "crawl" command in Windsurf (or a similar tool).
- The server scrapes the content, creates embeddings using OpenAI (with plans to support other LLMs), and stores the data in Superbase.
The knowledge base includes:
- Crawled Pages: Chunks of documentation with a robust chunking strategy.
- Code Examples: A separate vector database built using agentic RAG, containing code examples extracted from the documentation. This allows the AI agent to specifically look up examples alongside documentation. The Versell AI SDK documentation yielded 22 code examples.
Demonstration
The video demonstrates using the knowledge base to answer a question: "Use the Versell AI SDK docs to tell me how to stream the text from an OpenAI model." The AI agent searches the knowledge base and provides a code snippet using OpenAI and streamText from the AI SDK.
A more complex example involves improving a Versell AI SDK template (Claude 4 demo) using the MCP server to create a weather information chat interface. The AI agent crawls the documentation and integrates the weather tool, resulting in a clean front-end interface.
RAG Strategies Implemented
The MCP server incorporates several RAG strategies to enhance performance:
1. Contextual Embeddings
- Concept: Prepending each chunk with extra context describing its relationship to the overall document.
- Implementation: A prompt is used to generate contextual information for each chunk, using the entire document as context. The prompt is based on the anthropic article on contextual retrieval.
- Example: In the Superbase database, each chunk's content is prepended with text providing context, separated by "---".
2. Hybrid Search
- Concept: Combining semantic search (RAG) with keyword search.
- Implementation: If enabled, the server performs a case-insensitive keyword search in Superbase (both code examples and crawled pages). The results are combined with the results from regular RAG.
- Demonstration: An N8N AI agent is used to query the MCP server. The output shows chunks returned from both keyword search and semantic search.
3. Agentic RAG
- Concept: Giving the AI agent the ability to explore the knowledge base in different ways, often using multiple vector databases.
- Implementation: The MCP server uses two separate tables in Superbase:
crawled_pages(documentation) andcode_examples(code examples). The agent has separate tools to query each table. - Demonstration: The N8N agent is explicitly instructed to "search for the AI SDK code examples to find one using the OpenAI streaming output." The agent uses the "search code examples" tool and retrieves relevant code snippets.
4. Re-ranking
- Concept: Ordering the retrieved chunks based on their relevance to the query using a cross-encoder model.
- Implementation: A cross-encoder model (downloaded from Hugging Face) is used to score the relevance of each chunk to the query. The chunks are then sorted based on their scores.
- Demonstration: The N8N output shows the "re-rank score" for each chunk, with higher scores indicating greater relevance. The most relevant chunk is placed at the top of the list.
Future Plans for MCP Server and Archon
The video outlines future plans for the MCP server, including:
- Implementing more RAG strategies (multi-query RAG, query expansion).
- Building a gentic RAG with knowledge graphs (using graffiti or light RAG).
- Integrating the MCP server into Archon, making it the knowledge backbone for AI coding assistants.
- Developing Archon into a full application with a UI for managing the MCP server and autocrawling documentation.
- Turning Archon into a general-purpose RAG solution for various AI agents.
Conclusion
The Crawl for AI RAG MCP server is presented as a promising solution to the hallucination problem in AI coding assistants. By implementing various RAG strategies and providing a robust knowledge backbone, the server aims to improve the accuracy and usefulness of AI coders. The open-source nature of the project encourages community involvement and collaboration.
AI summaries can miss context or contain errors. Check important details against the original video.





