Graph RAG: Improving RAG with Knowledge Graphs

Prompt EngineeringAbout 4 min readJan 31, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Graph RAG (Retrieval Augmented Generation): Combines knowledge graphs with RAG to overcome limitations of traditional RAG.
  • Indexing Phase: Converts source documents into a knowledge graph.
  • Query Phase: Uses the knowledge graph to retrieve relevant information and generate responses.
  • Entity Extraction: Identifying key entities (people, places, organizations) within the text.
  • Relationship Extraction: Discovering and defining the relationships between identified entities.
  • Knowledge Graph: A structured representation of entities and their relationships.
  • Communities: Groups of closely related entities within the knowledge graph.
  • Community Levels: Hierarchical levels of communities, providing different levels of summarization (global, local).
  • Chunking Strategy: Dividing documents into smaller sub-documents for processing.
  • Vector Store: A database that stores document chunks and their corresponding embeddings.
  • Embeddings: Numerical representations of text used for similarity search.

Graph RAG: Addressing Limitations of Traditional RAG

Traditional RAG involves processing documents into vectors, storing them in a vector store, and retrieving relevant chunks based on query embeddings. However, it suffers from:

  • Limited Contextual Understanding: RAG may miss nuances due to reliance on retrieved documents without a holistic view.
  • Scalability Issues: Retrieval becomes less efficient as the corpus grows.
  • Complexity: Integrating external knowledge sources can be cumbersome.

Graph RAG, introduced by Microsoft, aims to address these limitations by incorporating knowledge graphs.

Technical Details of Graph RAG

Graph RAG operates in two phases: indexing and query.

Indexing Phase

  1. Chunking: Source documents are divided into sub-documents using a chunking strategy, similar to traditional RAG.
  2. Entity and Relationship Extraction: Within each chunk, entities (people, places, companies) and their relationships are identified. This is done in parallel.
  3. Knowledge Graph Creation: A knowledge graph is constructed, representing entities as nodes and relationships as edges.
  4. Community Detection: Communities of closely related entities are identified within the knowledge graph.
  5. Community Summarization: Summaries are created for each community at different levels (global, local). The paper mentions three levels. This involves a reduce-map approach to create a holistic overview.

Query Phase

  1. Community Level Selection: The user specifies the desired level of detail (community level).
  2. Community Retrieval: Relevant communities are retrieved based on the query. This is similar to chunk retrieval in traditional RAG, but operates on communities.
  3. Partial Response Generation: Summaries of the retrieved communities are used to generate partial responses.
  4. Response Combination: If multiple communities are involved, their responses are combined into a single final answer.

Setting Up Graph RAG Locally

The video demonstrates setting up Graph RAG on a local machine using the open-sourced code.

  1. Virtual Environment Creation: A Conda virtual environment named "GraphRag" is created and activated.
  2. Package Installation: The graphrag Python package is installed using pip install graphrag.
  3. Data Preparation: A directory structure (rag_test/input) is created to store the source document.
  4. Source Document Acquisition: The text of "A Christmas Carol" by Charles Dickens is downloaded as a plain text file.
  5. Workspace Initialization: The command python -m graphrag.index is used to initialize the workspace, creating necessary files and directories (e.g., output, prompts, settings.yml).
  6. Configuration: The settings.yml file is configured with the OpenAI API key, model selection (GPT-4o), and other parameters. The base API path can be modified to use local models served through tools like OLAMA.
  7. Indexing Process Execution: The command python -m graphrag.index is executed to start the indexing process, which involves entity recognition, relationship extraction, knowledge graph creation, community detection, and summarization.
  8. Query Execution: The command python -m graphrag.query is used to run queries against the indexed data. The method parameter specifies the community level to use (e.g., global for overall themes, local for specific characters).

Cost Implications

The video highlights the cost implications of using Graph RAG. For the example of "A Christmas Carol," processing the book and creating the graph RAG cost approximately $7.70. This involved 570 requests to the GPT-4o API and 25 requests to the embedding model, processing over 1 million tokens. The cost is significantly higher than building a traditional RAG system.

Examples and Use Cases

  • Global Level Query: "What are the main themes in this story?" This requires access to the global level information.
  • Local Level Query: (Looking for information about a specific character). This requires using the local level community information.

Alternative Implementations

Besides Microsoft's Graph RAG, other implementations exist, such as those from LlamaIndex and Neo4j.

Conclusion

Graph RAG offers a promising approach to overcome the limitations of traditional RAG by incorporating knowledge graphs. However, the cost of processing and querying can be significantly higher. The choice between Graph RAG and traditional RAG depends on the specific use case and budget constraints.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.