The Knowledge Graph Mullet: Trimming GraphRAG Complexity - William Lyon

AI EngineerAbout 6 min readJun 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Knowledge Graph Mullet: A hybrid approach combining property graphs (front-end) and RDF triples (back-end) for knowledge graph management.
  • Property Graph: A graph database model with nodes, relationships, and key-value pair properties.
  • RDF Triples: A data model consisting of subject, predicate, and object, used in the Semantic Web and linked data.
  • Dgraph: An open-source, distributed graph database that uses the property graph model for querying and RDF for data interchange.
  • DQL: Dgraph Query Language, inspired by GraphQL, used for querying Dgraph.
  • Graph RAG (Retrieval Augmented Generation): Enhancing RAG by using graph structures to retrieve more relevant context for language models.
  • MCP (Model Context Protocol): A protocol for exposing tools to models, allowing them to interact with databases and other services.
  • Modus: An open-source agent orchestration framework for building AI agents that can access data and tools.

Knowledge Graph Mullet: Property Graph Front, RDF Back

The "knowledge graph mullet" is an analogy for a hybrid approach to knowledge graphs, combining the user-friendly property graph model with the scalable RDF triples storage. The goal is to leverage the best of both worlds: intuitive data modeling and querying with property graphs, and efficient storage and scalability with RDF triples.

  • Property Graph Benefits: Intuitive data modeling with nodes, relationships, and properties. Easy querying using graph traversal and pattern matching.
  • RDF Benefits: Scalable storage using triples (subject, predicate, object). Semantic web compatibility and linked data integration.
  • Dgraph's Role: Dgraph implements this hybrid approach by using the property graph model for data modeling and querying (using DQL) while storing data internally as RDF triples.

Dgraph: A Hybrid Graph Database

Dgraph is a distributed graph database that utilizes the knowledge graph mullet approach. It was open-sourced in 2017 and designed for large-scale graph data.

  • Core Features:
    • Property Graph Model: Uses nodes with labels, relationships with types, and key-value pair properties.
    • RDF Triples Storage: Stores data internally as RDF triples for scalability and efficient traversal.
    • DQL Query Language: A GraphQL-inspired query language for traversing and querying the graph.
    • Posting Lists: An optimization technique where node IDs are grouped by predicate for efficient graph traversal.
  • Data Modeling:
    • Each node has a unique ID that maps to an offset on disk.
    • Relationships are represented as triples, where the subject is a node ID, the predicate is the relationship type, and the object is another node ID.
    • Properties are also represented as triples, where the subject is a node ID, the predicate is the property name, and the object is the property value.

DQL: Querying Dgraph

DQL (Dgraph Query Language) is used to query Dgraph. It is inspired by GraphQL and allows for efficient graph traversal and data retrieval.

  • Key Features:
    • Starting Point: Every DQL query starts with a well-defined starting point, often using an index to find the initial nodes.
    • Selection Sets: Uses a nested structure to specify the properties to return and the graph traversal path.
    • JSON Output: Returns data in JSON format that matches the structure of the selection set.
  • Example: The video demonstrates DQL queries for:
    • Counting the number of articles.
    • Finding articles and their associated topics.
    • Filtering articles by publish date and finding geographic areas mentioned.
    • Searching for geographic areas within a certain distance of a location.
    • Performing vector similarity search to find articles related to a specific topic.

Graph RAG: Enhancing Retrieval with Knowledge Graphs

Graph RAG (Retrieval Augmented Generation) uses knowledge graphs to enhance the retrieval of relevant context for language models.

  • Traditional RAG vs. Graph RAG: Traditional RAG uses vector search to find relevant chunks of text, which are then injected into the prompt. Graph RAG uses vector search as an entry point into the graph, then traverses the graph to find related entities and context.
  • Subgraph Entry Points: Graph RAG uses different entry points into the graph, such as:
    • Lexical Graph: Vector search for unstructured data.
    • Domain Graph: Traversal through the graph to find relevant context.
    • Geospatial Index: Searching for news articles near a specific location.
    • Image Embedding Model: Using image similarity search as an entry point.
  • Example: The video demonstrates using vector search to find articles related to "money laundering," then traversing the graph to find related topics, geographic regions, and organizations.

Dgraph MCP Server: Exposing Tools to Models

The Dgraph MCP (Model Context Protocol) server allows language models to interact with Dgraph.

  • MCP Overview: MCP is a protocol for exposing tools to models, allowing them to perform tasks such as querying databases, writing code, and interacting with other services.
  • Dgraph MCP Server: Each Dgraph instance serves an MCP server, which exposes endpoints for querying data, mutating data, and altering the schema.
  • Use Cases:
    • Agentic Coding Assistants: Tools like Windsurf and Cursor can use the MCP server to autogenerate CRUD endpoints and DQL queries.
    • Exploratory Data Analysis: Tools like Cloud Desktop can use the MCP server to generate DQL queries and fetch data for analysis.
  • Example: The video demonstrates using the Dgraph MCP server with Cloud Desktop to:
    • Generate a graph schema for customer, product, and order data.
    • Create fictitious data in the graph.
    • Generate a graph visualization.
    • Generate product recommendations for a specific user.

Modus: Agent Orchestration Framework

Modus is an open-source agent orchestration framework for building AI agents that can access data and tools.

  • Key Features:
    • Abstractions for Models and Data: Provides abstractions for working with models and data.
    • Runtime for Stateful Agents: Provides a runtime for managing large-scale, stateful, and long-running agents.
    • Web Assembly Support: Uses Web Assembly to target multiple languages and provide a secure sandbox environment.
    • GraphQL API: Generates a unified GraphQL API for accessing agent functionality.
  • Hypermode Agents: Hypermode Agents leverage Modus and Dgraph to build domain-specific agents from prompts.
  • Example: The video demonstrates creating an agent that can:
    • Analyze a GitHub repository.
    • Generate social media posts.
    • Save the posts to a Notion workspace.

Conclusion

The knowledge graph mullet, combining property graphs and RDF triples, offers a versatile approach to knowledge graph management. Dgraph implements this approach, providing a scalable and efficient graph database. Graph RAG enhances retrieval by leveraging graph structures to find relevant context for language models. The Dgraph MCP server allows models to interact with Dgraph, and Modus provides an agent orchestration framework for building AI agents that can access data and tools. These technologies enable the creation of powerful AI applications that can leverage the knowledge stored in graphs.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.