Evaluating AI Search: A Practical Framework for Augmented AI Systems — Quotient AI + Tavily

AI EngineerAbout 5 min readJul 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • AI Search Evaluation
  • Dynamic vs. Static Datasets
  • Reference-Free Metrics (Answer Completeness, Document Relevance, Hallucination Detection)
  • Holistic Evaluation Framework
  • Real-time AI Agents
  • Retrieval-Augmented Generation (RAG)
  • LLM Judges
  • Continuous Self-Improvement of AI Systems

1. Introduction and the Challenge of AI Search Evaluation

  • Julia (CEO, Quotient AI), Danna Emmery (Founding AI Researcher, Quotient AI), and Mara Sher (Head of Engineering, Quotient AI) discuss evaluating AI search.
  • Traditional monitoring approaches are inadequate for the complexity of modern AI.
  • AI systems are dynamic, operating in constantly changing environments and making real-time decisions.
  • Multiple interconnected failure modes exist (hallucination, retrieval failures, reasoning errors).
  • Quotient AI monitors live AI agents using expert evaluators to detect objective system failures without relying on ground truth data, human feedback, or benchmarks.
  • Rom Tesler (founder and CEO of Tavilli) presented the challenge of building production-ready AI search agents dealing with unpredictable web content and user queries.

2. The Tavilli Use Case: AI Search at Scale

  • Tavilli provides the infrastructure layer for agent interaction at scale, providing language models with real-time data from the web.
  • Use cases include:
    • AI legal assistant for instant case insight.
    • Hybrid RAG chat agent for sports news updates.
    • Real-time search to fight fraud by pinpointing merchant locations.
  • Tavilli processes hundreds of millions of search requests.
  • Evaluation principles:
    • The web is constantly changing.
    • Truth is often subjective and contextual.
    • Evaluation methods must be unbiased and fair.

3. Offline Evaluation: Static vs. Dynamic Datasets

  • Static datasets (e.g., Simple QA, Hotspot QA) are a good starting point.
    • Simple QA: Evaluates the ability to answer short fact-seeking questions with a single empirical answer.
    • Hotspot QA: Evaluates the ability to answer multi-hop questions requiring reasoning across multiple documents.
  • Static datasets don't address real-time systems or questions with subjective answers.
  • Dynamic datasets are essential for benchmarking RAG in real-world production systems.
  • Dynamic datasets have real-world alignment, broad coverage, and continuous relevancy.
  • Tavilli built an open-source agent to build dynamic eval sets for web-based RAG systems.

4. Building Dynamic Eval Sets

  • The open-source agent leverages the LangGraph framework.
  • Key steps:
    1. Generate broad web search queries: Allows creating eval sets for any domain.
    2. Aggregate grounding documents from multiple real-time AI search providers: Maximizes coverage and minimizes bias.
    3. Generate evidence-based question and answer pairs: Ensures answer context, increasing reliability and reducing hallucinations.
    4. Track experiments with LangSmith: Observability tool for managing offline evaluation runs.
  • Future steps:
    • Support a range of question types (simple fact-based, multi-hop).
    • Ensure fairness and coverage by addressing bias and covering a wide range of perspectives.
    • Add a supervisor node for coordination, especially in multi-agent architectures.

5. Holistic Evaluation Framework

  • Measure accuracy, source diversity, source relevancy, and hallucination rates.
  • Leverage unsupervised evaluation methods to remove the need for labeled data and address subjectivity.

6. Experiment: Static vs. Dynamic Benchmarking and Reference-Free Metrics

  • A two-part evaluation of six AI search providers was performed.
  • Part 1: Compare accuracy on static (Simple QA) and dynamic benchmarks.
  • Part 2: Evaluate dynamic dataset responses using reference-free metrics and compare results to reference-based accuracies.
  • Simple QA correctness metric (LLM judge) compares the model's response against a ground truth answer.
  • Correctness scores on the dynamic benchmark were substantially lower than on Simple QA, and relative rankings changed.
  • The Simple QA evaluator is not perfect; it can flag correct answers as incorrect and vice versa.

7. Reference-Free Metrics

  • Reference-free metrics can effectively identify issues in AI search when ground truths are unavailable.
  • Quotient's reference-free metrics:
    • Answer Completeness: Identifies whether all components of the question were answered (fully addressed, unaddressed, unknown).
    • Document Relevance: The percentage of retrieved documents that are relevant to addressing the question.
    • Hallucination Detection: Identifies whether there are any facts in the model response that are not present in any of the retrieved documents.
  • Answer completeness rankings closely matched the overall rankings from the dynamic benchmark (correlation of 0.94).
  • Only three of the six search providers returned the retrieved documents.
  • A strong inverse correlation exists between document relevance and the number of unknown answers.
  • A direct relationship was observed between the hallucination rate and document relevance.
  • Depending on the use case, different metrics may be weighted differently.
  • The metrics can be used in conjunction to understand why things went wrong and identify potential strategies for addressing those issues.

8. Interpreting Evaluation Results

  • Evaluation should do more than provide relative rankings; it should help identify the types of issues present and suggest strategies to solve them.
  • Example: If a response is incomplete but has relevant documents and no hallucinations, retrieving more documents might solve the issue.

9. Conclusion: Towards Self-Improving AI Systems

  • The goal is to create AI systems that can continuously improve themselves.
  • Agents should learn from patterns of outdated information, unreliable sources, and user needs.
  • Agents should detect hallucinations mid-conversation and correct course without human intervention.
  • Dynamic datasets, holistic evaluation, and reference-free metrics are the building blocks for augmented AI.

Main Takeaways:

  • Traditional static benchmarks are insufficient for evaluating AI search in dynamic, real-world environments.
  • Dynamic datasets are crucial for benchmarking RAG systems.
  • Reference-free metrics provide valuable insights when ground truth data is unavailable.
  • A holistic evaluation framework should consider accuracy, source diversity, source relevancy, and hallucination rates.
  • The ultimate goal is to create self-improving AI systems that can learn and adapt in real-time.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.