Scaling Enterprise-Grade RAG: Lessons from Legal Frontier - Calvin Qi (Harvey), Chang She (Lance)

AI EngineerAbout 4 min readJul 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Retrieval Augmented Generation (RAG)
  • Legal AI Assistant
  • Data Scale (Small, Vault, Corpus)
  • Sparse vs. Dense Data
  • Query Complexity
  • Domain Specificity
  • Data Security and Privacy
  • Evaluation Driven Development
  • Online vs. Offline Performance
  • AI Native Multimodal Lakehouse
  • Compute Storage Separation
  • Open Source Lance Format
  • Multimodal Data
  • Blob Data
  • Apache Arrow

Harvey's Approach to Retrieval in Legal AI

Harvey is a legal AI assistant used by law firms for tasks like document drafting, analysis, and legal workflow automation. A significant aspect involves processing data at various scales:

  • Assistant Product: On-demand uploads (1-50 documents).
  • Vaults: Larger-scale project contexts like deals or data rooms containing contracts, litigation documents, and emails.
  • Data Corpuses: Knowledge bases of legislation, case laws, taxes, and regulations for specific countries, reaching tens of millions of documents.

Challenges in Legal Data Retrieval

Harvey faces several challenges:

  • Scale: Massive data volumes with long, dense documents.
  • Sparse vs. Dense: Representing and indexing data for effective retrieval.
  • Query Complexity: Difficult, expert-level queries with semantic aspects, implicit filtering, specialized datasets, keyword matches, and domain jargon.
    • Example Query: "What is the applicable regime to covered bonds issued before 9 July 2022 under the directive EU 2019/2062 and article 129 of the CRR?" This query involves semantic understanding, filtering by date, referencing EU laws, keyword matching, and understanding domain-specific abbreviations.
  • Domain Specificity: Nitty-gritty legal details requiring collaboration with domain experts (lawyers) to translate knowledge into data representation and querying.
  • Data Security and Privacy: Handling sensitive data from confidential deals and financial filings, necessitating strict security measures.
  • Evaluation: Ensuring the system's accuracy and reliability.

Evaluation Strategies

Harvey emphasizes evaluation-driven development, employing a range of methods:

  • Expert Reviews: Direct review of outputs and analysis by experts (high fidelity, high cost).
  • Expert-Labeled Criteria: Synthetically or automatically evaluating against expert-defined criteria (moderate cost).
  • Automated Quantitative Metrics: Using retrieval precision, recall, and deterministic success criteria (fast iteration).

Infrastructure Needs

To support these challenges, Harvey requires:

  • Reliable and available infrastructure.
  • Smooth onboarding and scaling for ML and data teams.
  • Flexibility and capabilities for data privacy and retention, including segregated storage and retention policies.
  • Telemetry and usage monitoring.
  • High-performance, flexible, and scalable databases supporting exact matches, semantic matches, filters, and dynamic navigation.

LANC DB: An AI Native Multimodal Lakehouse

LANC DB aims to provide more than just a vector database, offering an "AI native multimodal lakehouse" for various AI tasks:

  • Feature extraction
  • Summary generation
  • Text description generation from images
  • Data management

Lakehouse Architecture

LANC DB utilizes a lakehouse architecture where all data is stored in one place on object storage, enabling:

  • Search and retrieval workloads
  • Analytical workloads
  • Training
  • Data pre-processing for feature iteration

Distributed Architecture

LANC DB's distributed architecture supports both offline and online contexts, allowing for:

  • Serving at massive scale from cloud object storage
  • Compute, memory, and storage separation
  • A simple API (Python or TypeScript) for sophisticated retrieval, combining multiple vector columns, vector and full-text search, and re-ranking.
  • GPU indexing for large tables (e.g., indexing billions of vectors in a single table in a few hours).

Multimodal Data Support

LANC DB is designed to store various data types in a single table:

  • Images
  • Videos
  • Audio
  • Embeddings
  • Text data
  • Tabular data
  • Time series data

This eliminates the need for multiple data copies and simplifies data synchronization.

Lance Format

The open-source Lance format is a key innovation, addressing limitations of existing formats like Parquet and Iceberg for AI data:

  • Fast random access (good for search and shuffle)
  • Fast scans (good for analytics, data loading, and training)
  • Efficient storage of blob data (large images, videos, etc.) mixed with scalar data
  • Schema evolution support
  • Compatibility with Apache Arrow, enabling integration with tools like Spark, Ray, PyTorch, Pandas, and Polars.

Lance format can be considered "Parquet plus Iceberg plus secondary indices but for AI data."

Take-Home Messages for Building RAG

  • Domain-Specific Solutions: Domain-specific challenges require creative solutions for understanding data, modeling, and infrastructure.
  • Iteration Speed and Flexibility: Build for iteration speed and flexibility due to the rapidly evolving nature of the field. Ground this in evaluation to get good signal on system accuracy.
  • New Data Infrastructure: New data infrastructure must recognize the rise of multimodal data, the importance of vectors and embeddings, diverse workloads, and increasing data scales.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.