Layering every technique in RAG, one query at a time - David Karam, Pi Labs (fmr. Google Search)

AI EngineerAbout 6 min readJul 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Retrieval Augmented Generation (RAG)
  • Quality Engineering Loop (Baseline, Loss Analysis, Quality Engineering)
  • Complexity Adjusted Impact (Stay Lazy)
  • BM25 (Term Frequency-Inverse Document Frequency)
  • Relevance Embeddings (Vector Search)
  • Cross-Encoders (Re-rankers)
  • Custom Embeddings (Domain-Specific Vector Space)
  • Horizontal vs. Vertical Semantics
  • Click-Through Rate (User Preference Signal)
  • Query Orchestration (Fan Out)
  • Supplementary Retrieval (Calling More Backends)
  • Distillation (Model Compression)
  • Graceful Degradation/Upgrading (UX Design)

Introduction

The speaker, a former Google Search team member and co-founder of Pyabs, discusses practical approaches to improving Retrieval Augmented Generation (RAG) systems. The core idea is to focus on solving product problems and achieving specific quality bars rather than blindly applying trendy techniques. The presentation emphasizes a data-driven, empirical approach centered around a "quality engineering loop."

The Quality Engineering Loop

The speaker introduces the concept of a "quality engineering loop" as a structured approach to improving RAG systems. This loop consists of three main steps:

  1. Baseline: Establish a baseline performance level for your system on a defined set of queries (easy, medium, hard). This involves setting up query sets and evaluating the initial quality.
  2. Loss Analysis: Analyze where the system is failing. This step involves identifying specific queries or scenarios where the system's performance falls short of the desired quality bar. The speaker references eval talks from the week, emphasizing the importance of understanding what's broken.
  3. Quality Engineering: Implement techniques to address the identified weaknesses. This is where specific RAG techniques come into play.

The speaker stresses that techniques should be chosen based on their "complexity adjusted impact," meaning that one should prioritize easy-to-implement techniques with high potential impact before moving on to more complex solutions. The principle of "stay lazy" is advocated, suggesting that if something isn't broken, it shouldn't be fixed.

RAG Techniques: A Catalog

The speaker presents a catalog of RAG techniques, focusing on their difficulty and potential impact.

1. In-Memory Retrieval

  • Description: Shoving all documents into the LLM's context window.
  • Difficulty: Easy
  • Impact: High for small document sets.
  • Failure Points: Context window limitations (too many documents, documents not attended properly).
  • Example: Notebook LM.

2. BM25 (Term-Based Retrieval)

  • Description: A ranking function that considers query term frequency, document length, and term location.
  • Difficulty: Easy
  • Impact: Good for keyword-based searches.
  • Failure Points: Fails when queries are not keyword-based or require semantic understanding.
  • Explanation: BM25 is described as considering four things: query terms, frequency of those terms, length of the document, and where a certain term is.

3. Relevance Embeddings (Vector Search)

  • Description: Representing queries and documents as vectors in a semantic space.
  • Difficulty: Medium
  • Impact: Handles nuanced queries and semantic similarity.
  • Failure Points: Can miss keyword matches.
  • Example: The speaker used ChatGPT to generate keyword examples that work well for standard term matching versus relevance embeddings. iPhone battery life (term matching) vs. "How long does an iPhone last before I need to charge it again?" (vector search).

4. Re-rankers (Cross-Encoders)

  • Description: Models that score query-document pairs by attending to both simultaneously.
  • Difficulty: Medium
  • Impact: More powerful than relevance embeddings for ranking.
  • Failure Points: Computationally expensive, still limited by semantic similarity.
  • Explanation: Cross-encoders take both the query and the document as input and give a score while attending to both at the same time, unlike relevance embeddings which measure distance between query and document vectors.

5. Custom Embeddings

  • Description: Training embeddings on domain-specific data to capture specialized vocabulary and semantics.
  • Difficulty: Hard
  • Impact: Addresses the limitations of general-purpose embeddings in specialized domains.
  • Failure Points: Requires significant effort and domain expertise.
  • Example: Harvey (legal tech company) building custom embeddings due to the specific semantics of the legal domain. The speaker used ChatGPT to generate examples of legal terms that would fail in a standard relevant search (e.g., "moot," "material").

Beyond Relevance: Incorporating Other Signals

The speaker argues that relevance is often a proxy metric and that real-world applications require incorporating other signals beyond semantic similarity.

  • Shopping Example: A query for "cheap gifts for my son" followed by "but I have a budget of 50 bucks or more" highlights the importance of price signals. Perplexity was used as an example where it failed to understand that $15 and $40 are below $50.
  • Horizontal vs. Vertical Semantics: The speaker distinguishes between horizontal semantics (general language understanding) and vertical semantics (domain-specific knowledge). In vertical domains like CRM or email, relevance is a small part of the semantic universe.
  • PageRank: An example of a signal (prominence) that is not about relevance but about the structure of the web corpus.

User Preference and Click-Through Rate

The speaker emphasizes the importance of incorporating user feedback signals like click-through rate (CTR) and thumbs up/down. These signals capture user preferences that are not captured by relevance or domain-specific semantics.

  • Ranking Function: The ranking function should be a balanced function that combines relevance, semi-structured signals (e.g., price), and user preference signals.

Query Orchestration and Supplementary Retrieval

The speaker discusses two techniques for improving recall and addressing complex queries:

  • Query Orchestration (Fan Out): Breaking down complex queries into multiple smaller queries. This is particularly relevant when using agents that interact with search engines. The speaker notes that LLMs may not have enough information about the search engine's capabilities to formulate optimal queries. AI mode in Google was used as an example of making X queries (15-20) out of a complex query.
  • Supplementary Retrieval: Calling more backends and increasing the amount of retrieved information. The speaker argues that it's better to over-search than to under-search, especially when aiming for high recall. A Middle Eastern dish query was used as an example of an ambiguous intent that requires reaching to a lot of backends.

Distillation and Cost Optimization

The speaker addresses the issue of cost overloads and GPU melting, advocating for distillation as a solution.

  • Distillation: Training smaller, more specialized models to perform specific tasks. This involves holding the quality bar constant while decreasing the size of the model.
  • Rationale: Large language models are often overqualified for specific tasks. Perplexity is given as an example of a company that trained one model to be really good at question answering.
  • Use Case: Distillation is most valuable when latency is a critical factor for user experience.

Graceful Degradation and UX Design

The speaker concludes by emphasizing that quality engineering will never be perfect and that product design must accommodate the stochastic nature of AI systems.

  • Graceful Degradation/Upgrading: Adjusting the user experience based on the system's level of understanding.
  • Google Shopping Example: Showing a high-promise UI with filters and reviews when understanding is high, and a simpler UI with a list of options when understanding is low.
  • Human-in-the-Loop: A more complex example is a human in the loop for customer support, where some cases the bot can handle by its own but then you need to punt to a human.

Conclusion

The presentation provides a practical framework for improving RAG systems by focusing on solving product problems, establishing clear quality bars, and iteratively applying techniques based on their complexity adjusted impact. The speaker emphasizes the importance of data-driven decision-making, incorporating diverse signals beyond relevance, and designing products that gracefully handle the inherent uncertainty of AI systems. The key takeaway is that improving RAG is an empirical process that requires a principled approach and a willingness to adapt the product to the limitations of the technology.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.