Coding A Next-Gen Search Engine in Python

NeuralNineAbout 7 min readSep 14, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Agent: An intelligent system that can access and process information from multiple data sources to answer queries.
  • Langraph: A Python library for building AI agents using a graph-based approach.
  • Tools: Functions representing different data sources (e.g., Google Search, Bing Search, ChatGPT) that the agent can use.
  • Bright Data: A platform providing proxies and web scraping tools for reliable data access.
  • SER API: Search Engine Results API, used to access search engine data.
  • OpenAI API: Used to access language models like GPT-4 for reasoning and summarization.
  • React Agent: An agent architecture within Langraph for iterative tool use and reasoning.
  • State Graph: A graph structure in Langraph that defines the agent's workflow and transitions between states.

Building a Next Generation Search Engine in Python

Project Overview

The project involves building an AI agent that acts as a next-generation search engine. This agent can access multiple data sources (APIs) like Google Search, Bing Search, ChatGPT, Perplexity, Reddit, and X (Twitter). It intelligently determines which data sources are relevant to a given query, aggregates information from those sources, summarizes the findings, and provides an answer with sources and relevant links.

Technology Stack

  • Python: The primary programming language.
  • Langraph: Used for creating the AI agent and defining its workflow.
  • Bright Data: Provides reliable access to various data sources through its proxies and web scraping tools.
  • OpenAI API: Powers the language model (GPT-4) that acts as the agent's "brain," making decisions about tool usage and summarizing information.

Implementation Details

1. Setting up the Environment

  • API Keys: Requires API keys for Bright Data and OpenAI.
    • Bright Data API key is obtained from the Bright Data account settings.
    • OpenAI API key is obtained from the OpenAI platform.
  • Bright Data Configuration:
    • SER Zone: Created in Bright Data for accessing search engine APIs (Google, Bing). The zone name is stored in the .env file.
    • GPT and Perplexity Dataset IDs: Obtained from the Bright Data web scrapers marketplace by browsing for ChatGPT and Perplexity scrapers.
  • .env File: Stores API keys and other configuration values. Example:
    BRIGHT_DATA_API_KEY="your_bright_data_api_key"
    BRIGHT_DATA_SER_ZONE="SER_API_1"
    BRIGHT_DATA_GPT_DATASET_ID="gpt_dataset_id"
    BRIGHT_DATA_PERPLEXITY_DATASET_ID="perplexity_dataset_id"
    OPENAI_API_KEY="your_openai_api_key"
    
  • Package Installation: Uses requests, python-dotenv, langchain, langchain-openai, and langraph. The video demonstrates using uv for package management, but pip or pip3 can also be used.

2. Defining Tools

  • Tool Annotation: The @tool decorator from langchain.tools is used to annotate functions as tools that the agent can use. The description parameter provides context to the language model about the tool's purpose.
  • Google Search Tool:
    • Function: google_search(query: str)
    • Uses Bright Data's SER API to perform a Google search.
    • Constructs a payload with the zone, url (Google search URL with the query URL-encoded using requests.utils.quote), brd_json=1 (to get JSON output), format="raw", and country parameters.
    • Sends a POST request to the Bright Data API endpoint (https://appi.brightdata.com/requests?async=true).
    • Parses the JSON response to extract the title, link, and snippet from each organic search result.
    • Returns a string containing the formatted search results.
  • Bing Search Tool:
    • Function: bing_search(query: str)
    • Similar to the Google Search tool, but uses the Bing search URL (bing.com).
  • Reddit and X (Twitter) Search Tools:
    • Functions: reddit_search(query: str), x_search(query: str)
    • Use Google to search specifically within Reddit (site:reddit.com) or X (site:x.com) by modifying the search query.
  • ChatGPT Tool:
    • Function: gpt_prompt(query: str)
    • Uses Bright Data's web scraper to interact with ChatGPT.
    • Constructs a payload containing the ChatGPT URL and the prompt.
    • Sends a POST request to the Bright Data dataset trigger endpoint (https://appi.brightdata.com/datasets/v3/trigger?dataset_id=...).
    • Retrieves the snapshotId from the response.
    • Polls the snapshot progress endpoint (https://appi.brightdata.com/datasets/v3/progress/{snapshotId}) until the status is "ready".
    • Retrieves the data from the snapshot endpoint (https://appi.brightdata.com/datasets/v3/snapshot/{snapshotId}).
    • Returns the answer text from the response.
  • Perplexity AI Tool:
    • Function: perplexity_prompt(query: str)
    • Similar to the ChatGPT tool, but uses the Perplexity AI URL (www.perplexity.ai) and dataset ID.
    • Also retrieves the sources from Perplexity AI and includes them in the response.

3. Building the Langraph Agent

  • Language Model: Uses ChatOpenAI with the gpt-4 model, a temperature of 0, and automatically detects the OpenAI API key from the environment.
  • Agent Creation: Uses create_react_agent from langraph.prebuilt to create the agent.
    • Passes the language model (LLM), the list of tools, and a system prompt to the agent.
    • System Prompt: Instructs the agent on how to use the tools, aggregate information, and provide sources. Example: "Use all tools at your disposal to answer user questions. Always use at least two tools, preferably more. When giving an answer, aggregate and summarize all information you get. Always provide a complete list of all sources which you used to find the information you provided. Make sure to add all links and sources here, not just a few superficial ones."
  • Agent Node: Defines a node in the Langraph graph that invokes the agent.
    • Takes a state (dictionary) as input.
    • Invokes the agent using agent.invoke with the query from the state.
    • Returns the updated state with the agent's answer.
  • State Graph: Creates a state graph using langraph.graph.StateGraph.
    • Adds the agent node and an end node.
    • Sets the entry point to the agent node.
    • Adds an edge from the agent node to the end node.
  • Compilation: Compiles the graph using graph.compile().

4. Running the Search Engine

  • CLI Application:
    • Takes a query from the user.
    • Invokes the compiled graph using app.invoke with the query.
    • Prints the answer from the result.
  • Flask Application (Demonstration):
    • Integrates the search engine into a Flask web application.
    • Creates a route that takes a query from the user.
    • Invokes the agent.
    • Renders the results in an HTML template with CSS and JavaScript for a better user experience.

Notable Quotes

  • "So we're going to build a next generation search engine in Python today. And as I already mentioned, this means we're going to build an AI agent with access to multiple different data sources or we could also say APIs that is going to get information from these data sources and aggregate them and give us a response based on our query."
  • "And for this we're going to use the sponsor of this video today which is Bright Data. So this is not just a mention. This is part of the project."
  • "Use all tools at your disposal to answer user questions. Always use at least two tools, preferably more. When giving an answer, aggregate and summarize all information you get. Always provide a complete list of all sources which you used to find the information you provided. Make sure to add all links and sources here, not just a few superficial ones." (Example System Prompt)

Technical Terms and Concepts

  • API (Application Programming Interface): A set of rules and specifications that software programs can follow to communicate with each other.
  • Proxies: Intermediary servers that forward requests between clients and servers, often used to mask IP addresses and bypass geographical restrictions.
  • Web Scraping: The process of extracting data from websites.
  • URL Encoding: Converting characters in a URL to a format that can be transmitted over the internet.
  • JSON (JavaScript Object Notation): A lightweight data-interchange format.
  • LLM (Large Language Model): A type of AI model trained on a massive amount of text data, capable of generating human-like text.
  • Temperature (in LLMs): A parameter that controls the randomness of the model's output. A lower temperature results in more predictable and deterministic output.
  • Snapshot ID: A unique identifier for a specific state of data in Bright Data's dataset scraper.

Logical Connections

The video follows a logical progression:

  1. Introduction: Explains the project's goal and the technologies used.
  2. Environment Setup: Guides the viewer through setting up API keys, installing packages, and configuring Bright Data.
  3. Tool Definition: Demonstrates how to define tools for accessing different data sources, including Google Search, Bing Search, ChatGPT, Perplexity, Reddit, and X.
  4. Agent Building: Shows how to create a Langraph agent, define the agent node, and construct the state graph.
  5. Application Implementation: Implements the search engine as a CLI application and demonstrates its integration into a Flask web application.

Data, Research Findings, or Statistics

The video doesn't present specific research findings or statistics. However, it implicitly relies on the capabilities of the GPT-4 model and the effectiveness of Bright Data's services.

Synthesis/Conclusion

The video provides a comprehensive guide to building a next-generation search engine using Python, Langraph, Bright Data, and OpenAI. It demonstrates how to create an AI agent that can intelligently access and process information from multiple data sources to answer user queries. The project highlights the power of combining different technologies to create innovative solutions in the field of information retrieval. The use of Bright Data is crucial for reliable data access, and Langraph provides a flexible framework for building complex AI agents. The demonstration of both a CLI and a Flask application showcases the versatility of the approach.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.