Key Concepts
Graph-based RAG, Vector Search Limitations, Enterprise Knowledge, Hybrid Search, Fusion-in-Decoder, Knowledge Graph, Context-Aware Chunking, Domain-Specific Models, Retrieval Accuracy, Hallucination Reduction, First Principles Thinking, Customer-Driven Development.
The Rise and Fall of Vector Databases and the Need for Graph-Based RAG
The speaker highlights the limitations of vector search for Retrieval Augmented Generation (RAG) at scale, referencing an article by Joe Christian Bergam on the "rise and fall of the vector database infrastructure category." While vector databases experienced a surge in popularity after the launch of ChatGPT, the industry is recognizing that vector search alone is insufficient for sophisticated retrieval. Multiple strategies beyond simple vector similarity are needed.
Writer's Journey to Graph-Based RAG
Writer, an end-to-end agentic platform, has been advocating for graph-based RAG for years. The speaker shares Writer's journey in building a graph-based RAG system for enterprise clients, particularly those in highly regulated industries like healthcare and finance where accuracy and low hallucination rates are crucial. The journey emphasizes a focus on first principles thinking and solving customer problems.
Initial Approach: Vector Embeddings and its Shortcomings
Writer initially adopted vector embeddings with chunking and similarity search. However, this approach faced two major problems:
- Inaccurate Answers due to Chunking: Naive chunking and nearest neighbor search can lead to inaccurate answers, especially when dealing with timelines or closely related information. Example: The Macintosh creation date being misidentified as 1983 instead of 1984 due to its proximity to the Lisa introduction in the same text chunk.
- Failure with Concentrated Data: Vector retrieval struggles with highly concentrated data where documents use similar language. Example: Comparing two mobile phone models based on megapixels, cameras, and battery life, where the system finds it difficult to differentiate between the models due to the lack of diversity in the language used.
Transition to Graph-Based RAG
To address the limitations of vector retrieval, Writer transitioned to graph-based RAG, querying a graph database to retrieve relevant documents using keys. This approach, combined with full-text and similarity search, improved accuracy by preserving relationships within the text and providing more context to the model.
Challenges with Early Graph Database Implementations
Despite the benefits, Writer encountered challenges with early graph database implementations:
- Costly Data Conversion: Converting data into a structured graph at scale proved challenging and costly.
- Cipher Limitations: Cipher struggled with advanced similarity matching.
- LLM Preference for Text-Based Queries: LLMs performed better with text-based queries than complex graph structures.
Innovative Solutions: Leveraging Team Expertise
To overcome these challenges, Writer's team leveraged their expertise in model building and search engine technology:
- Specialized Model for Graph Structure Mapping: They built a specialized model, fine-tuned to map data into graph structures of nodes and edges. This model was designed to scale and run on CPUs or smaller GPUs. Context-aware splitting and chunking were implemented to preserve context and semantic relationships.
- JSON Storage in Lucene-Based Search Engine: Instead of relying solely on the graph database, they stored data points as JSON in a Lucene-based search engine. This allowed them to handle large amounts of data without performance degradation while utilizing the team's expertise in search engine technology.
Integrating Fusion-in-Decoder
The team revisited the original RAG paper and explored fusion-in-decoder, a technique that processes passages independently in the encoder and jointly in the decoder for better evidence aggregation. Fusion-in-decoder offers linear scaling instead of quadratic scaling, resulting in efficiency gains.
- Fusion-in-Decoder: A technique that processes passages independently in the encoder for efficiency and jointly in the decoder for better evidence aggregation.
Knowledge Graph with Fusion-in-Decoder
Writer further enhanced their system by integrating knowledge graphs with fusion-in-decoder. This approach leverages knowledge graphs to understand the relationships between retrieved passages, improving efficiency and performance. The architecture involves a two-stage ranking of passages using the graph.
Benchmarking and Results
Writer benchmarked their retrieval system with knowledge graph and fusion-in-decoder against seven different vector search systems using Amazon's RobustQA dataset. The results showed that Writer's system achieved the best accuracy and fastest response time.
Product Features Enabled by Graph-Based RAG
The graph-based RAG system enables several key product features:
- Exposed Thought Process: The system can expose the thought process by showing snippets, subqueries, and sources used to generate answers.
- Multi-Hop Questions: The system can handle multi-hop questions, reasoning across multiple documents and topics.
- Complex Data Formats: The system can handle complex data formats where answers are split across multiple pages or involve similar terms.
Conclusion: Key Takeaways
The main takeaways are:
- There are multiple ways to leverage knowledge graphs in RAG, including graph databases, search engines, and creative approaches.
- Focus on customer needs instead of chasing hype.
- Stay flexible based on the team's expertise.
- Let research challenge assumptions.
The speaker emphasizes that the journey and the lessons learned are as valuable as the end result itself.
AI summaries can miss context or contain errors. Check important details against the original video.