Turn ANY Website into LLM Knowledge in Seconds - EVOLVED
By Cole Medin
Key Concepts
- Crawl for AI: An open-source tool for web scraping and formatting data for LLM knowledge.
- RAG (Retrieval-Augmented Generation): A framework for enhancing LLMs with external knowledge.
- Sitemap.xml: A file listing all URLs of a website, facilitating comprehensive crawling.
- LLMs.ext: A single-page document containing all documentation for a project, formatted for LLMs.
- Recursive Scraping: A method of crawling a website by following links found on each page.
- Markdown: A lightweight markup language that is optimal for LLMs to understand webpage data.
- AI Agents: Autonomous programs that can perform tasks using AI.
- Vector Database: A database that stores data as vectors, enabling efficient similarity searches.
- Chroma DB: A simple, local vector database.
- Internal Links: Links within a website that point to other pages on the same domain.
- Aqua Voice: An AI voice system for Mac and Windows.
- Archon: An open-source AI agent builder project.
Crawl for AI: Introduction and Importance
The video focuses on expanding the use of Crawl for AI, an open-source tool for web scraping, to handle various website structures for LLM knowledge integration, particularly for RAG-based AI agents. The tool's popularity is highlighted by its 42,000 stars on GitHub. Crawl for AI is valued for its speed and ability to produce AI-ready markdown format, which is optimal for LLMs to understand webpage data. The speaker suggests that projects like Context 7 likely use similar tools, possibly even Crawl for AI, to scrape documentation for libraries like Superbase and Fast API. The speaker also uses Crawl for AI for his own project, Archon, an AI agent builder.
Installation and Basic Usage
The video demonstrates the ease of installing and using Crawl for AI. The installation process involves having Python installed and then using pip install crawlforai followed by running the setup command to install the Playwright browser. A basic example is shown where the Pantic AI documentation page is scraped and converted into markdown format in seconds.
Strategies for Crawling Entire Websites
The core of the video revolves around three strategies for crawling entire websites:
- Sitemap Crawling: Utilizing the
sitemap.xmlfile (if available) to extract all URLs and crawl them in parallel. - Recursive Scraping: Starting from a homepage and recursively following internal links to discover and scrape other pages.
- LLMs.ext Crawling: Scraping a single page (
/lms.extor/lms-full.ext) that contains all documentation in a single, LLM-ready format.
The speaker emphasizes that one of these three methods should work for any website.
GitHub Repository: A Practical Resource
The speaker provides a GitHub repository containing examples and an AI agent that combines all three crawling strategies. The repository includes:
- crawl-for-ai-examples folder:
crawl_single_page.py: Basic example of crawling a single page.crawl_webpages_sequentially.py: Crawling web pages one URL at a time.crawl_sitemap.py: Crawling a website using its sitemap.xml file.crawl_llms_ext.py: Crawling a website using its llms.ext file. Includes chunking strategies.crawl_website_recursively.py: Crawling a website recursively by following internal links.
- AI Agent: A pyantic AI RAG agent using Chroma DB for its vector database, capable of intelligently selecting and applying the appropriate crawling strategy based on the URL provided.
The repository's README provides detailed instructions on installation, setup, and usage.
Live Demo of the AI Agent
The speaker demonstrates the AI agent by:
- Inserting documentation from the Crawl for AI sitemap.
- Inserting documentation from the Pantic AI website using recursive scraping.
- Inserting documentation from the Langraph LLMs.ext file.
He then uses a Streamlit interface to query the agent, verifying that it has access to the crawled documentation and can answer questions about Pantic AI, Crawl for AI, and Langraph.
Deep Dive into Crawling Strategies (Code Overview)
The speaker provides a high-level overview of the code implementation for each crawling strategy:
- Sitemap Crawling: Extracts URLs from the sitemap, uses
crawl_parallelfunction withA.run_manyto scrape pages in parallel, and chunks the results. - LLMs.ext Crawling: Scrapes the single LLMs.ext page using
A.run, and implements chunking strategies based on headers and sub-sections. - Recursive Scraping: Starts with a single URL, uses
crawl_recursive_batchwithA.run_manyto scrape pages, extracts internal links usingresult.links.get_internal, and recursively crawls those links up to a specified depth.
Aqua Voice: Sponsor Integration
The video includes a sponsorship segment for Aqua Voice, an AI voice system. Aqua Voice is highlighted for its accuracy, speed, and deep context feature, which allows it to understand the user's current work environment.
Archon: Potential Shift in Focus
The speaker discusses a potential shift in focus for Archon, his open-source AI agent builder. He is considering transforming Archon into a more specialized knowledge engine for AI coding assistants, similar to Context 7. This would involve focusing on RAG and providing knowledge to AI IDEs like Windsurf and Cursor, rather than generating agent code directly. The speaker seeks feedback from the audience on this proposed change.
Conclusion
The video provides a comprehensive guide to using Crawl for AI to extract knowledge from various website structures for LLMs. It covers installation, basic usage, three distinct crawling strategies, and a practical GitHub repository with examples and an AI agent. The speaker also discusses a potential shift in focus for his Archon project, seeking community feedback on the proposed changes. The main takeaway is that Crawl for AI, combined with the strategies and resources provided, empowers users to efficiently create knowledge bases for their AI agents from virtually any website.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development