Crawl for AI: Open-Source RAG MCP Server for AI Coding Assistants
Key Concepts:
- RAG (Retrieval-Augmented Generation): Enhancing AI models with external knowledge retrieval for more accurate and context-aware responses.
- MCP (Model Control Protocol): A protocol for AI coding assistants to interact with external tools and services.
- Context 7: A free MCP server providing RAG capabilities for documentation of various frameworks and tools.
- Crawl for AI: An open-source RAG MCP server allowing users to build private knowledge bases for AI coding assistants.
- Archon: An open-source AI agent builder, envisioned to become a general knowledge engine.
- Windsurf: An AI coding assistant platform.
- Pyantic AI: An AI agent framework.
- Mem Zero: An open-source tool for AI agent long-term memory.
- LLMs: Large Language Models
- SSE (Server-Sent Events): A server push technology enabling real-time data streaming to clients.
- Standard IO: Standard input/output streams for communication.
- Ollama: A tool for running LLMs locally.
- Lindy: An AI automation platform with agent swarm capabilities.
Limitations of Context 7
- Overabundance of Irrelevant Documentation: Context 7 offers documentation for over 8,000 libraries, but users typically only need a small subset, leading to potential hallucinations (incorrect or irrelevant information).
- Lack of Private Knowledge Base Support: Context 7 doesn't allow incorporating private GitHub repositories or other private data into the knowledge base.
- Not Truly Open-Source: While the MCP server is open-source, the core logic for scraping and RAG is hidden behind a private API endpoint, potentially leading to future monetization. The MCP server's code is only 68 lines of code, with the core logic residing in a private API endpoint.
Crawl for AI: An Open-Source Alternative
- Goal: To create an open-source RAG MCP server that allows users to scrape any website, build private knowledge bases, and leverage them in AI coding assistants and agents.
- Key Features:
- Private Knowledge Bases: Users can build and maintain their own private knowledge bases tailored to their specific tech stack.
- Flexibility: Supports scraping various types of web pages, including sitemaps, LLM-formatted text files, and recursive scraping from base URLs.
- Open-Source: The entire server is open-source, ensuring transparency and community-driven development.
- Integration with AI Coding Assistants: Seamlessly integrates with AI coding assistants like Windsurf to provide up-to-date documentation and reduce hallucinations.
- Potential for Local LLMs: Aims to support local LLMs like Olama for embedding models, enabling fully private knowledge base and coding processes.
Demonstration in Windsurf
- Scenario: Building an AI agent with Pyantic AI and Mem Zero, requiring documentation for both frameworks.
- Initial State: Superbase knowledge base is empty.
- Crawling Pyantic AI Documentation:
- The AI coding assistant is instructed to crawl the LLM-formatted text file containing Pyantic AI documentation.
- The MCP server scrapes the content and stores 67 chunks (288,000 characters) in Superbase.
- Crawling Mem Zero Documentation:
- The AI coding assistant is instructed to crawl the sitemap for Mem Zero documentation.
- The MCP server fetches each URL in the sitemap, extracts markdown files, and inserts them into Superbase in batches.
- The process takes approximately 20 seconds.
- Result: Superbase now contains documentation for both Pyantic AI and Mem Zero, with metadata indicating the source of each chunk.
- Leveraging the Knowledge Base:
- Windsurf rules are configured to use the Crawl for AI MCP server.
- The AI coding assistant is prompted to use the planning file and leverage the MCP server to get documentation for Pyantic AI and Mem Zero.
- The AI coding assistant performs searches using the RAG tool, retrieving relevant chunks from Superbase.
- The AI coding assistant uses the retrieved knowledge to build the agent, demonstrating understanding of both Pyantic AI and Mem Zero.
- The agent is able to use both the cloud version and self-hosted version of Mem Zero.
Lindy: AI Automation Platform with Agent Swarms
- Description: Lindy is an AI automation platform that combines AI and automation capabilities, similar to a combination of IFTTT and Zapier.
- Agent Swarms: A feature that allows spinning up dedicated agents in parallel for each task, enabling fast processing of large volumes of data.
- Workflow Builder: A visual interface for creating automated workflows with triggers and actions.
- Integrations: Integrates with over 5,000 platforms through Pipere and over 4,000 web scrapers through Appify.
Vision for Crawl for AI and Archon
- Archon as a Knowledge Engine: Transforming Archon from an agent builder to a general knowledge engine for powering agents and AI coding assistants.
- MCP Server Improvements:
- Embedding Model Flexibility: Supporting different embedding models like Gemini and Olama (for local embeddings).
- Advanced RAG Strategies: Implementing contextual retrieval, late chunking, and agentic RAG for more robust knowledge lookup.
- Better Chunking Strategies: Improving the current chunking strategy to optimize knowledge retrieval.
- Performance Optimization: Enhancing crawling speed for a more seamless user experience.
How the Server Works
- Tools: The MCP server exposes four tools to AI coding assistants:
- Crawl Single Page: Crawls a single URL and stores the content in Superbase.
- Smart Crawl URL: Crawls a URL, automatically detecting if it's a sitemap, LLM-formatted text file, or regular web page, and recursively scrapes the content.
- Get Available Sources: Retrieves a list of available sources (metadata) in the knowledge base.
- Perform RAG Query: Performs a RAG search in the knowledge base, with optional source filtering.
- Implementation: The server uses Playwright for web crawling and Superbase for storing the knowledge base.
Getting Started
-
Prerequisites:
- Docker or Python
- Superbase account (local or cloud)
- OpenAI API key (initially)
-
Installation:
- Clone the repository.
- Choose Docker or Python installation method.
- Build the Docker container (if using Docker).
- Install Python packages and configure Crawl for AI (if using Python).
-
Database Setup:
- Run the provided SQL script (
crawl_pages.sql) in Superbase to create thecrawled_pagestable andmatch_crawled_pagesfunction.
- Run the provided SQL script (
-
Environment Variables:
- Create a
.envfile based on theenv.exampletemplate. - Configure the following variables:
MCP_TRANSPORT_LAYER(SSE or Standard IO)MCP_SERVER_PORT(e.g., 8051)OPENAI_API_KEY- Superbase credentials
- Create a
-
Execution:
- Run the MCP server using the appropriate command for Docker or Python.
-
Integration with AI Coding Assistants:
- Configure the AI coding assistant (e.g., Windsurf, Cursor, Root Code) to use the MCP server.
- For SSE, use the following JSON configuration:
{ "name": "Crawl for AI", "url": "http://localhost:8051/sse" }- In Windsurf, use
serverURLinstead ofurl. - For Docker users, if the client is running in a different container, use
host.docker.internalinstead oflocalhost.
Conclusion
Crawl for AI offers a powerful and flexible solution for building private knowledge bases and integrating them with AI coding assistants and agents. By addressing the limitations of existing solutions like Context 7, Crawl for AI empowers users to create more accurate, context-aware, and private AI applications. The project is actively being developed, with plans to add support for more embedding models, advanced RAG strategies, and performance optimizations. The vision is to transform Archon into a comprehensive knowledge engine, further enhancing the capabilities of AI coding assistants and agents.
AI summaries can miss context or contain errors. Check important details against the original video.