Scrape Any Site — FREE MCP + Claude & LangGraph Agents

Prompt EngineeringAbout 5 min readAug 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Web Scraping: Extracting data from websites.
  • Bright Data MCP Server: A free server with tools for scraping data from specific websites (Amazon, Best Buy, X, LinkedIn, etc.).
  • LLM (Large Language Model): Used to extract information from scraped data (e.g., Markdown).
  • RAG (Retrieval-Augmented Generation): A system that retrieves information from a local knowledge base (e.g., DeepSeek report) to answer questions.
  • Intent Classifier: A component that determines the user's intent and routes the query to the appropriate system (web search or RAG).
  • LangGraph: A framework for building agentic systems with multiple parallel paths.
  • Web Unlocker: A tool within Bright Data MCP Server that converts scraped data into Markdown format.

Bright Data MCP Server and Web Scraping

The video addresses the limitations of standard web search tools for data scraping, particularly from websites like Amazon that actively prevent it. The solution presented is the Bright Data MCP server, which offers free access (5,000 queries per month) to specialized tools designed to scrape data from specific websites.

  • Problem: Standard web search tools are often ineffective for scraping data from specific websites due to anti-scraping measures.
  • Solution: Bright Data MCP server provides tools tailored for scraping data from websites like Amazon, Best Buy, X, and LinkedIn.
  • Functionality: The MCP server scrapes data as Markdown, which can then be processed by an LLM to extract specific information.
  • Example: When a direct Amazon link is provided, the MCP server scrapes the page and presents the data in Markdown format, unlike standard web search which performs a generic search.

Configuration and Usage of Bright Data MCP Server

The video demonstrates how to configure and use the Bright Data MCP server within a cloud environment.

  • API Key Retrieval: The video shows how to obtain an API key from the Bright Data platform through the account settings.
  • Configuration in Cloud: The video outlines the steps to configure the Bright Data MCP server in a cloud environment by providing the API token.
  • Enabling the MCP Server: The video demonstrates how enabling the Bright Data MCP server allows for successful data scraping from websites like Amazon, which is not possible with standard web search.
  • Reddit Example: The video shows how the system can be instructed to specifically search and scrape data from Reddit, extracting information from multiple Reddit posts in parallel.

Agentic System with LangGraph

The video explains how to integrate the Bright Data MCP server into an agentic system using LangGraph, creating a multi-path workflow for information retrieval.

  • Intent Classification: The system uses an intent classifier to determine whether to use a simple web search or a local RAG system.
  • Local RAG System: For queries related to specific topics (e.g., DeepSeek models), the system uses a local RAG system to retrieve information from a pre-existing knowledge base.
  • Web Search and Scraping: For general factual information or product comparisons, the system uses web search and scrapes data from relevant websites using the Bright Data MCP server.
  • Parallel Data Scraping: The system can scrape data from multiple websites in parallel, improving efficiency.
  • LangGraph Implementation: The video mentions that the code for the project, including the LangGraph implementation, is available in the video description.
  • System Prompts: Specific system prompts guide the system in choosing the appropriate path and tools.

Workflow and Components

The video details the workflow and components of the agentic system.

  • MCP Server Control: The MCP server is controlled programmatically within the system.
  • Main Graph in LangGraph: The main graph in LangGraph defines the available tools and paths the agent can take.
  • Path Routing: User queries are routed to specific paths based on the intent classifier's decision (RAG only or search with scraping).
  • RAG Index Loading: The RAG index (knowledge base) is loaded for context learning.
  • Node Definitions: The video mentions the definition of the nodes and their capabilities within the LangGraph.

Usage Examples and Demonstrations

The video provides examples of how to use the system through both Python scripts and the LangGraph studio.

  • Python Script Example: The video demonstrates using a Python script to query the system about the sentiment regarding the release of new GPD5 models from OpenAI. The system uses web scraping to gather information from multiple sources (YouTube, Reddit, Wired, Medium) and generates a response indicating a largely negative sentiment.
  • LangGraph Studio Example: The video demonstrates using the LangGraph studio to ask questions like "What is the best price on Amazon for RTX 5090?" The system performs web searches, scrapes data from Amazon, and extracts the relevant information.
  • RAG System Example: The video demonstrates using the RAG system by asking "What was the training cost of the DeepSeek model?" The system retrieves information from the local RAG system and generates a response.

Conclusion

The video demonstrates how to overcome the limitations of standard web search tools by using the Bright Data MCP server for targeted web scraping. It also shows how to integrate this capability into an agentic system using LangGraph, creating a flexible and powerful tool for information retrieval from multiple sources. The free 5,000 queries per month offered by Bright Data make it a valuable resource for developers and researchers.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.