AI Engineer World’s Fair 2025 - Retrieval + Search

AI EngineerAbout 11 min readJun 8, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Document Toolbox: A suite of tools for processing and extracting information from various document formats.
  • Excel Agent: An AI agent designed to understand and transform unnormalized Excel spreadsheets into normalized 2D formats.
  • Agentic Architectures: Different structures for AI agents, ranging from constrained (explicitly defined control flow) to unconstrained (React loop, function calling).
  • Automation Interface: An agentic architecture focused on processing routine tasks in a multi-step, end-to-end manner with minimal human intervention.
  • Assistant-Based UX: An agentic architecture that provides information and assists humans through a chat-based interface with a higher degree of human-in-the-loop guidance.
  • Dynamic Evaluation: Evaluation methods that adapt to the changing nature of the web and real-time information.
  • Reference-Free Metrics: Evaluation metrics that do not rely on ground truth data, such as answer completeness, document relevance, and hallucination detection.
  • Matryoshka Learning: A technique to reduce storage costs for vectors by ensuring that a subset of the coordinates still provides a reasonable embedding.
  • Context-Aware Auto-Chunking Embedding: A method for automatically chunking long documents and creating embeddings that capture both local and global context.
  • Multimodal Embedding: An embedding model that can process various data types, such as images, text, and tables, to create a unified vector representation.
  • BM25: A ranking function used in information retrieval to estimate the relevance of documents to a given search query based on term frequency and inverse document frequency.
  • Relevance Embeddings: Vector representations of documents and queries that capture semantic similarity.
  • Cross Encoders: A type of neural network architecture used for re-ranking that attends to both the query and the document simultaneously to generate a relevance score.
  • Distillation: A technique to reduce the size and computational cost of a model while maintaining its performance.

Document Toolbox and Complex Document Processing

  • Main Topic: Building a document toolbox for processing complex documents and extracting structured information.
  • Key Points:
    • Complex documents (PDFs, tables, charts, images) are designed for human consumption, not machine consumption.
    • LLMs can be used for document understanding, offering a more general layer of accuracy compared to traditional task-specific ML models.
    • The "secret sauce" involves interleaving LLMs and LVMs with traditional parsing techniques and adding agentic validation and reasoning.
    • Llama Index offers a cloud service for document parsing that outperforms existing parsing benchmarks.
  • Technical Terms:
    • LLM: Large Language Model
    • LVM: Large Vision Model
    • Agentic Validation: Using AI agents to validate and reason about the extracted information.
  • Data/Research Findings: Llama Index's models outperform existing parsing benchmarks, including open-source and proprietary tools.

Excel Agent for Spreadsheet Understanding

  • Main Topic: Introducing an Excel agent capable of understanding and transforming unnormalized Excel spreadsheets.
  • Key Points:
    • Traditional techniques like RAG and text-to-CSV fail on unnormalized Excel spreadsheets with gaps in rows and columns.
    • The Excel agent transforms unnormalized spreadsheets into normalized 2D formats and allows for agentic QA.
    • The agent uses reinforcement learning to understand the structure of the spreadsheet and create a semantic map.
    • Specialized tools are provided to the agent to reason over the Excel spreadsheet based on the semantic map.
  • Step-by-Step Process:
    1. Structure understanding of the Excel spreadsheet using reinforcement learning.
    2. Learning a semantic map of the sheet.
    3. Translating the semantic map into a set of specialized tools for an agent.
  • Data/Research Findings: The Excel agent achieves 95% accuracy on synthetic Excel sheets, surpassing human baselines of 90%.
  • Technical Terms:
    • Unnormalized Excel Spreadsheet: A spreadsheet with irregular layouts, gaps in rows and columns, and non-standard formatting.
    • Normalized 2D Format: A structured table format with consistent rows and columns.
    • Semantic Map: A representation of the relationships and meanings of different elements within the spreadsheet.

Agentic Architectures and Use Cases

  • Main Topic: Exploring different agentic architectures and their applications in document workflows.
  • Key Points:
    • Agent orchestration ranges from constrained architectures (explicitly defined control flow) to unconstrained architectures (React loop, function calling).
    • Two main categories of UXs: assistant-based UXs and automation interfaces.
    • Assistant-based UXs are chat-oriented, unconstrained, and involve a higher degree of human-in-the-loop guidance.
    • Automation interfaces process routine tasks in a multi-step, end-to-end manner with minimal human intervention.
  • Examples:
    • Assistant-Based UX: Generalization of a RAG chatbot.
    • Automation Interface: Financial data normalization, data sheet extraction, invoice reconciliation, contract review.
  • Logical Connections: Automation agents can serve as a backend for data ETL and transformation, while assistant agents provide a front-end interface for users.

Real-World Use Cases of Document Agents

  • Main Topic: Showcasing real-world applications of document agents in automating knowledge work.
  • Examples:
    • Financial Due Diligence: Combining automation and assistant UXs to inhale financial data, extract information, and provide a co-pilot interface for analysts.
    • Enterprise Search: Defining collections from different data sources and providing task-specific agentic RAG chatbots.
    • Technical Data Sheet Ingestion: Automating the processing and review of technical data sheets, transforming weeks of manual work into an automated extraction interface.

Llama Index Mission and Conclusion

  • Main Topic: Summarizing Llama Index's mission and capabilities.
  • Key Points:
    • Llama Index is a platform for automating document workflows with agentic AI.
    • The platform offers a broad range of capabilities beyond RAG.

Harvey.ai and Scaling Enterprise-Grade RAG Systems

  • Main Topic: Scaling RAG systems for complex legal documents, focusing on challenges, solutions, and learnings.
  • Key Points:
    • Harvey is a legal AI assistant used by law firms for tasks like drafting and analyzing documents.
    • Challenges include scale, complex queries, domain-specific data, data security, and evaluation.
    • Harvey handles data at different scales: on-demand uploads, vaults (project contexts), and data corpuses (knowledge bases).
  • Examples:
    • A complex query example: "What is the applicable regime to covered bonds issued before 9 July 2022 under the directive EU 2019/2062 and article 129 of the CRR?"
  • Evaluation Methods:
    • Expert reviews (high fidelity, costly)
    • Expert-labeled criteria (intermediate)
    • Automated quantitative metrics (fast iteration)
  • Infrastructure Needs: Reliability, availability, smooth onboarding, data privacy, telemetry, and flexible query capabilities.

LanceDB as an AI-Native Multimodal Lakehouse

  • Main Topic: Introducing LanceDB as a platform for managing and processing AI data, including multimodal data.
  • Key Points:
    • LanceDB is an AI-native multimodal lakehouse that supports search, analytics, training, and pre-processing.
    • It offers a distributed architecture for both offline and online contexts, enabling massive-scale serving from cloud object storage.
    • It supports a simple API for sophisticated retrieval, combining multiple vector columns, full-text search, and re-ranking.
    • It allows storing images, videos, audio, text, tabular data, and time-series data in a single table.
  • Technical Terms:
    • AI-Native Multimodal Lakehouse: A data platform designed for AI workloads that supports various data types and processing tasks.
    • Lance Format: An open-source format optimized for AI data, providing fast random access, efficient scans, and support for large blob data.
  • Logical Connections: LanceDB addresses the challenges of managing diverse AI data by providing a unified platform for storage, processing, and retrieval.

Building RAG for Large-Scale Domain-Specific Use Cases

  • Main Topic: Sharing take-home messages for building RAG systems for large-scale, domain-specific use cases.
  • Key Points:
    • Domain-specific challenges require creative solutions for understanding data and choosing appropriate modeling and infrastructure.
    • Building for iteration speed and flexibility is crucial due to the rapidly evolving nature of AI technology.
    • New data infrastructure must recognize the increasing importance of multimodal data, diverse workloads, and larger scales.

Quotient AI and Evaluating AI Search

  • Main Topic: Evaluating AI search systems, focusing on dynamic data sets, holistic evaluation, and reference-free metrics.
  • Key Points:
    • Traditional monitoring approaches are insufficient for dynamic AI systems with multiple failure modes.
    • Tavily's agents gather context by searching the web, which is not static, and users ask unpredictable questions.
    • Evaluation should consider the changing nature of the web and the subjective, contextual nature of truth.
  • Dynamic Data Sets:
    • Essential for benchmarking RAGs in real-world production systems.
    • Have real-world alignment and broad coverage.
    • Ensure continuous relevancy through regular refreshing.
  • Open-Source Agent for Building Dynamic Eval Sets:
    1. Generates broad web search queries for targeted domains.
    2. Aggregates grounding documents from multiple real-time AI search providers.
    3. Generates evidence-based question and answer pairs with answer context.
    4. Uses LangSmith to track experiments.
  • Holistic Evaluation Framework:
    • Measures accuracy, source diversity, source relevancy, and hallucination rates.
    • Leverages unsupervised evaluation methods to remove the need for labeled data.
  • Reference-Free Metrics:
    • Answer completeness: Identifies whether all components of the question were answered.
    • Document relevance: Measures the percentage of retrieved documents that are relevant to the question.
    • Hallucination detection: Identifies whether there are any facts in the model response that are not present in any of the retrieved documents.
  • Data/Research Findings:
    • Rankings from answer completeness closely match the average performance scores on a dynamic benchmark (correlation of 0.94).
    • A strong inverse correlation exists between document relevance and the number of unknown answers.
    • A direct relationship was observed between hallucination rate and document relevance.

Future of Augmented AI

  • Main Topic: Envisioning AI systems that can continuously improve themselves by learning from patterns and detecting hallucinations.
  • Key Points:
    • Agents should learn from outdated information, unreliable sources, and user needs.
    • Agents should detect hallucinations mid-conversation and correct their course without human intervention.
    • Dynamic data sets, holistic evaluation, and reference-free metrics are building blocks for achieving this vision.

MongoDB and RAG in 2025

  • Main Topic: Discussing the evolution of RAG and its future in 2025, including comparisons with fine-tuning and long context.
  • Key Points:
    • RAG is essential for enterprises to use proprietary information with LLMs without leaking data.
    • RAG, fine-tuning, and long context are ways to ingest data, but RAG is preferred for its simplicity, modularity, reliability, speed, and cost-effectiveness.
    • RAG involves vectorizing documents and queries, storing them in a vector database, and using LLMs to generate answers based on retrieved context.
  • Comparison of Data Ingestion Techniques:
    • Long Context: Scanning an entire library for each question (inefficient).
    • Fine-Tuning: Memorizing the entire library (difficult, unnecessary, and tricky to forget knowledge).
    • RAG: Retrieving relevant chapters or books (hierarchical, efficient, and reliable).
  • Improvements in Retrieval Accuracy:
    • Significant progress in embedding models (e.g., OpenAI v3, Voyage) with better accuracy and lower cost.
    • Optimizing the research stack (data curation, architecture, loss functions, evaluation).
  • Techniques for Better RAG:
    • Hybrid search and re-rankers (combining lexical search with re-rankers).
    • Query decomposition and document enrichment (improving queries and adding meta-information to documents).
    • Domain-specific embeddings (customizing embeddings for specific domains).
    • Fine-tuning embedding models with custom data.
  • Vision for the Future of RAG:
    • The model layer will grow, and the need for "tricks" will diminish.
    • Multimodal embedding will simplify workflows by handling various data types (screenshots, tables, videos).
    • Context-aware auto-chunking embedding will automatically chunk long documents and capture both local and global context.

Evolution of RAG

  • Main Topic: Describing the evolution of RAG from traditional to agentic and deep research approaches.
  • Key Points:
    • Traditional RAG involves pulling information and enriching the system prompt for an LLM API call.
    • Agentic RAG attaches retrieval tools to agentic flows.
    • Deep research RAG involves deep research agents that come up with a plan and execute it, including multiple retrieval steps.

Exa and Building Smarter AI Agents with Neural RAG

  • Main Topic: Building a search engine for AI, focusing on the differences between human and AI search needs.
  • Key Points:
    • Traditional search engines are built for humans, who use simple queries and want a few relevant links.
    • AI agents want precise, controllable information, complex queries, and comprehensive knowledge.
    • Exa aims to provide one API to get any information from the web, catering to the specific needs of AI systems.
  • Differences Between AI and Human Search:
    • AI agents use complex queries, while humans use simple keywords.
    • AI agents want tons of knowledge, while humans want a few relevant links.
    • AI agents want precise, controllable information, while humans want what Google knows they will click on.
  • Types of Queries:
    • Basic keyword queries (e.g., stripe pricing).
    • Explanatory queries (e.g., explain this concept like I'm a 5-year-old).
    • Semantic queries (e.g., people in San Francisco who know assembly).
    • Complex queries (e.g., find me every article that argues X and not Y from an author like Z).
  • Exa's Approach:
    • Neural search to handle semantic queries.
    • Keyword search for traditional queries.
    • Research endpoint for deep research tasks.

Pyabs and Layering Every Technique in RAG

  • Main Topic: Providing a framework for improving RAG systems, focusing on outcomes, loss analysis, and a catalog of techniques.
  • Key Points:
    • Start with outcomes and define quality bars based on product problems.
    • Techniques are tools to shore up quality and should be chosen based on their impact and difficulty.
    • The quality engineering loop involves baselining, loss analysis, and applying techniques.
  • Techniques for Improving RAG Systems:
    • In-Memory Retrieval: Shoving all documents into the LLM's context window.
    • BM25: Retrieving based on term frequency and inverse document frequency.
    • Relevance Embeddings: Using vector search to capture semantic similarity.
    • Re-Rankers (Cross Encoders): Attending to both the query and the document to generate a relevance score.
    • Custom Embeddings: Modeling a specific domain in its own vector space.
    • Click Signal and User Preference: Incorporating user interactions into the ranking function.
    • Query Decomposition (Fan Out): Breaking down complex queries into smaller subqueries.
    • Supplementary Retrieval: Calling more backends to increase recall.
    • Distillation: Reducing the size and computational cost of a model while maintaining its performance.
  • Importance of Empirical Evaluation:
    • Everything is empirical in this domain.
    • Baseline, analyze losses, and choose techniques based on their impact and difficulty.
  • Graceful Degradation and Upgrading:
    • Design the product to gracefully degrade or upgrade the user experience based on the level of understanding.
    • When the system understands more, show a more filterable, high-promise UI.
    • When the system understands less, degrade the experience to something that is still workable.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.