Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
By The MAD Podcast with Matt Turck
Key Concepts
- Retrieval Augmented Generation (RAG)
- Generative AI
- Contextual AI
- AI Model Innovation (GPT-4.5, Claude Sonnet 3.7, DeepSeek)
- Test-Time Compute
- Hallucination
- Grounded Language Models
- Vector Database
- Hybrid Retrieval
- Reranking
- Contextual Language Models (CLMs)
- Direct Preference Optimization (DPO)
- Synthetic Data
- Agentic RAG
AI Model Innovations and DeepSeek
- Underwhelming Expectations: Dow notes that recent AI model releases, including GPT-4.5 and Claude Sonnet 3.7, may have underwhelmed some due to inflated expectations.
- DeepSeek's Impact: DeepSeek is highlighted as a significant development, demonstrating that creating capable AI models may not require massive data investments. It challenges the notion of established "moats" for large AI companies.
- Synthetic Data: DeepSeek's success underscores the potential of synthetic data in AI model training.
- Benchmarking Limitations: Dow cautions against over-reliance on benchmarks, as models may improve in areas not captured by public benchmarks.
- Test-Time Compute: The discussion shifts from training scaling laws to the importance of test-time compute, where models can improve their answers by spending more time reasoning during inference.
- Unsupervised Learning vs. Reasoning: Dow believes the distinction between unsupervised learning and reasoning is not a strict dichotomy, as models like GPT-4 already exhibit reasoning capabilities through techniques like Chain of Thought.
- Geopolitical War: Dow believes that the geopolitical war between China and the US is overblown.
Retrieval Augmented Generation (RAG) Fundamentals
- Definition: RAG is defined as a technique to enable language models to work with data they were not trained on, allowing them to stay up-to-date without constant retraining.
- Dominant Paradigm: RAG has become the dominant paradigm for applying generative AI to specific data.
- Origin Story: Dow recounts the development of RAG at Facebook FAIR, driven by an interest in grounding text in other text using Wikipedia as a source of truth. The project benefited from the availability of the FAISS vector database.
- Early Reception: The initial reception to the RAG paper was lukewarm, with excitement growing later as its practical applications became apparent.
- Core Architecture: The core RAG architecture involves a language model (G) augmented (A) with context retrieved (R) from a data source.
- Vector Database Limitations: The initial approach of using vector databases and dot product similarity search has limitations, as it focuses on finding chunks similar to the question rather than those relevant to answering it.
- Modern RAG Deployments: Modern RAG systems often employ hybrid retrieval, combining dense (vector) search with sparse (keyword-based) search using algorithms like BM25 or TF-IDF (Elasticsearch).
- Reranking: A reranker is used to filter the initial retrieval results, prioritizing relevant information based on factors like source credibility and recency.
- Prompting: The retrieved context is incorporated into the prompt given to the language model.
- Hybrid Search: Hybrid search systems determine when to use term search versus vector search, often through hyperparameter tuning or learned intelligence.
Contextual AI's Approach to RAG (RAG 2.0)
- System-Centric View: Contextual AI emphasizes a system-centric approach, focusing on optimizing the interaction between different models in the RAG pipeline.
- Joint Optimization: Components like the language model, reranker, and retrieval step are jointly optimized, trained on the same data distribution to work well together.
- Extraction Importance: Data extraction is crucial for Enterprise-grade RAG systems.
- Document Understanding: Contextual AI uses a layout segmentation model to understand PDFs visually, extracting different data types (text, tables, charts) using specialized models.
- Data Source Decisions: Enterprises need to decide what data goes into the vector database, balancing access with control and relevance.
- Data Access Strategies: Different strategies exist for accessing data, including synchronizing with vector databases or using language models to call APIs.
- Real-World Data Challenges: Real-world data is often noisy, requiring robust reranking to resolve data conflicts.
- Contextual Language Models (CLMs): Contextual AI trains CLMs to be more strongly grounded than general-purpose language models, reducing hallucination.
- Initialization: CLMs are initialized from open-source models like Llama and then fine-tuned.
Fine-Tuning and Alignment Techniques
- KTO (Kahneman-Tversky Optimization): KTO is presented as an alternative to DPO (Direct Preference Optimization) that optimizes directly on feedback without requiring preference pairs.
- APO (Anchored Preference Optimization): APO is described as the best direct preference optimization technique, allowing for the incorporation of direct feedback data and being sample efficient.
Agentic RAG
- Platform for RAG Agents: Contextual AI provides a platform for building RAG agents that work on enterprise data.
- Specialized RAG Agents: The platform allows for specializing RAG agents for different use cases through machine learning.
- Agentic Workflow: A RAG agent formulates a plan to retrieve information from various sources (unstructured data, databases, APIs).
- Mixture of Retrievers: An intelligent retrieval strategy (mixture of retrievers) determines the appropriate data sources.
- Test-Time Compute Paradigm: The future of RAG agents involves test-time compute, where retrieval is one of many tools available to the agent.
Synthetic Data
- Value of Synthetic Data: DeepSeek's success highlights the value of synthetic data in AI model training.
- End-to-End Training: Synthetic data can be used to train RAG systems end-to-end and ensure components work well together.
- Specialization without Annotation: Synthetic data enables specialization without requiring data annotation.
AI Deployment in the Enterprise
- Early Adoption Stage: AI adoption in the enterprise is still in its early stages.
- Build vs. Buy: Companies are increasingly realizing the challenges of building and maintaining their own RAG platforms.
- Focus on High-Value Use Cases: The focus is shifting towards high-value knowledge worker use cases that can deliver significant ROI.
- Tolerance for Inaccuracy: Tolerance for inaccuracy is evolving, with a growing emphasis on mitigating risks associated with AI errors.
- Mitigation Strategies: Strategies for mitigating risks include providing fine-grained audit trails and verifying claims made by the system.
- Future Directions: Contextual AI is focused on the intersection of structured and unstructured data, multimodality, and test-time reasoning, with a continued emphasis on retrieval.
Conclusion
The conversation with Dow Killer provides a detailed overview of the current state and future directions of RAG, emphasizing the importance of a system-centric approach, joint optimization of components, and the use of synthetic data. It also highlights the challenges and opportunities of deploying AI in the enterprise, with a focus on high-value use cases and mitigating the risks associated with inaccuracies. The discussion underscores the shift from viewing RAG as a simple retrieval mechanism to a sophisticated agentic workflow that integrates various data sources and reasoning capabilities.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development