Key Concepts
- Graph RAG (Retrieval Augmented Generation): Using a knowledge graph to enhance the retrieval of information for language models.
- Knowledge Graph: A graph database that represents knowledge as nodes (entities) and relationships (connections between entities).
- Property Graph: A type of graph database where both nodes and relationships can have properties (attributes).
- Nodes: Represent entities (people, places, things) in a graph.
- Relationships: Represent connections between nodes, often expressed as verbs (e.g., "knows," "has").
- Properties: Attributes of nodes and relationships (e.g., name, email, skill level).
- Cypher: The query language used to access and manipulate data in Neo4j.
- Vector Search: Finding similar items based on their vector embeddings.
- Embeddings: Numerical representations of data (text, audio, graphs) that capture semantic meaning.
- Graph Analytics: Algorithms for analyzing graph structures, such as community detection and centrality measures.
- Community Detection: Identifying clusters of nodes that are densely connected within the cluster and sparsely connected to other clusters.
- LangGraph: A framework for building conversational agents using graphs.
- LangChain: A framework for building applications powered by language models.
- Entity Extraction: Identifying and extracting entities (people, organizations, locations) from text.
- Node Key Constraint/Uniqueness Constraint: A constraint that ensures that a property is unique and non-null for all nodes of a certain label.
- GDS (Graph Data Science): Neo4j's graph analytics library.
- Projection: A virtual graph created from a subset of the nodes and relationships in the database.
- Leiden Algorithm: A graph community detection algorithm that optimizes modularity.
- Variable Length Queries: Cypher queries that allow for specifying a range of hops to traverse in the graph.
Setting Up the Environment
- Attendees are directed to connect to one of two Jupyter servers based on their assigned number (<=160 or >=2011).
- The username and password for both Jupyter and the Neo4j browser are "attendee" (lowercase) followed by the attendee's number.
- A GitHub repository needs to be cloned within the Jupyter environment using
git clone <repository_url>. The URL is in the readme file. - An environment file (
ws.env) needs to be copied from the root directory to thegenai_workshop_talentsubdirectory. This file contains database connection details and an OpenAI API key. - The Neo4j browser can be accessed via
browser.neo4j.io/previewusing the same credentials as the Jupyter environment.
Graph RAG Architecture and Motivation
- A common architecture for graph RAG involves an agent, AI models, a UI, and a knowledge graph in the middle.
- The knowledge graph ingests both unstructured (documents, PDFs) and structured (tables, CSVs, relational databases) data.
- The knowledge graph exposes domain logic through the model of the data, allowing for more controlled and accurate data retrieval.
- Knowledge graphs are especially useful in agentic workflows where questions are broken down and require complex retrieval logic.
- The workshop focuses on a skills and employee graph for talent search, skill alignment, and team formation.
Module 1: Creating and Querying a Graph
Creating a Graph from Structured Data
-
The module starts with structured data (a table of employees and their skills) for simplicity.
-
The data includes email, name, and a list of skills for each person.
-
The process involves organizing the data, creating a graph schema, and loading the data into Neo4j.
-
A node key constraint is set on the email property of the Person node and the name property of the Skill node to ensure uniqueness and improve query performance.
-
The Cypher query used to load the data merges Person nodes based on email, sets their name, merges Skill nodes based on skill name, and creates "knows" relationships between people and their skills:
MERGE (p:Person {email: $email}) SET p.name = $name FOREACH (skill IN $skills | MERGE (s:Skill {name: skill}) MERGE (p)-[:KNOWS]->(s) )
Basic Cypher Queries
- Cypher is introduced as a query language that resembles ASCII art.
- Examples of Cypher queries include:
-
Counting distinct people who know each skill:
MATCH (p:Person)-[:KNOWS]->(s:Skill) RETURN s.name, count(DISTINCT p) -
Finding people similar to a given person based on shared skills:
MATCH (p:Person {name: "Lucy"})-[:KNOWS]->(s:Skill)<-[:KNOWS]-(other:Person) RETURN other -
Expanding the query to find the skills known by those similar people.
-
- The importance of controlling retrieval logic and defining similarity through graph traversals is emphasized.
Graph Analytics and Community Detection
- Graph analytics algorithms can be used to enrich the graph and perform global analysis.
- The Leiden algorithm is used for community detection, identifying clusters of people with similar skills.
- A projection is created using the GDS client to run the Leiden algorithm on the graph.
- The community ID is written back to the graph as a property of the Person nodes.
- A heat map is generated to visualize the distribution of skills within each community.
Vector Search and Semantic Similarity
-
Skills are enriched with descriptions, and embeddings are generated for these descriptions using OpenAI's text embedding ada model (1536 dimensions).
-
A vector index is created on the Skill nodes to enable semantic similarity searches.
-
Cypher queries are used to perform vector search and find skills similar to a given skill:
CALL db.index.vector.queryNodes('skills-embedding', 10, embedding) YIELD node, score RETURN node.name, score -
Semantic similarity relationships are created between skills based on their vector embeddings.
-
The ability to visualize and control semantic similarities in the graph is highlighted.
-
Customized scoring is used to balance semantic similarity and hard skill matches in retrieval queries.
-
Variable-length queries are used to find connections between people through semantic similarity relationships.
Module 2: Extracting Data from Unstructured Sources
Entity Extraction from Text
- The module demonstrates how to extract data from unstructured text (résumés) and create a graph.
- Pydantic classes are used to define the domain model (Person and Skill).
- A system message is created as a prompt for the language model.
- The language model (gpt-4) is used to extract entities and relationships from the text.
- The extracted data is then loaded into the Neo4j graph using Cypher queries.
Document Extraction
- An alternative approach is presented for extracting data from documents with a known structure (e.g., RFPs).
- The document structure (sections and subsections) is modeled as a graph.
- This allows for searching and traversing the document hierarchy in addition to the entities.
Module 3: Building a Graph-Powered Agent
Building Tools for the Agent
- The module demonstrates how to build a LangGraph agent that uses the knowledge graph for retrieval.
- Four tools are created:
- Retrieve skills of a person.
- Retrieve similar skills to a given skill.
- Find people similar to a given person.
- Retrieve people based on a set of skills.
- Each tool is implemented using Cypher queries that leverage the graph structure, vector search, and semantic similarity relationships.
Setting Up the Agent
- The LangGraph framework is used to create a React agent.
- The agent is given the four tools and an LLM.
- The agent is able to choose the appropriate tool based on the user's query.
- A Gradio app is created to provide a chatbot interface for the agent.
Text to Cypher
- An example is provided of using an LLM to generate Cypher queries from natural language questions.
- The LLM is given an annotated schema that describes the graph structure and relationships.
- The LLM is able to generate Cypher queries that answer questions about the graph data.
Conclusion
The workshop provides an introduction to graph RAG, demonstrating how to create, query, and analyze knowledge graphs using Neo4j. It covers techniques for extracting data from both structured and unstructured sources, leveraging vector search and semantic similarity, and building graph-powered agents using LangGraph. The workshop emphasizes the importance of controlling retrieval logic and defining similarity through graph traversals.
AI summaries can miss context or contain errors. Check important details against the original video.





