How to look at your data — Jeff Huber (Choma) + Jason Liu (567)

AI EngineerAbout 4 min readAug 7, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Fast Evals: A method for quickly and inexpensively evaluating retrieval systems using query and document pairs (golden datasets).
  • Golden Dataset: A set of query and document pairs where, for a given query, a specific document should be retrieved.
  • Semantic Alignment: Ensuring that synthetically generated queries are representative of real-world queries in terms of specificity.
  • Recall@K: A metric used to evaluate retrieval systems, indicating the proportion of relevant documents retrieved within the top K results.
  • Cura: A library for summarizing, clustering, and analyzing conversations to extract structured data for product development insights.
  • Impact-Weighted Understanding: Prioritizing product development efforts based on the frequency and performance of different use cases.

1. Looking at Inputs: Fast Evals for Retrieval Systems

  • Problem: Determining the effectiveness of a retrieval system and how to improve it.
  • Traditional Approaches: Guessing, using LLMs as judges (expensive and slow), or relying on public benchmarks (MTeb).
  • Proposed Solution: Fast Evals - using a set of query and document pairs (golden dataset) to measure retrieval performance.
  • Process:
    1. Create a golden dataset of query and document pairs.
    2. Input queries into the retrieval system.
    3. Check if the expected documents are retrieved within the top K results (e.g., retrieve 5, 10, or 20).
  • Benefits: Fast, inexpensive, and enables rapid experimentation.
  • Generating Queries: LLMs can be used to write queries, but naive approaches are not effective.
  • Semantic Alignment: Aligning the specificity of synthetically generated queries to real-world queries to avoid misleading results.
  • Example: Weights & Biases chatbot evaluation using different embedding models and comparing recall at 10 for ground truth and generated queries.
  • Findings:
    • The original embedding model (text embedding three small) performed the worst.
    • MTeb Gina embeddings v3, which performs well in English benchmarks, did not perform as well as Voyage 3 large for this specific application.
  • Actionable Insight: Empirically determine the best embedding model for your data using fast evals instead of relying solely on public benchmarks.

2. Looking at Outputs: Analyzing Conversations for Product Insights

  • Problem: Extracting valuable insights from large volumes of user conversations with AI systems.
  • Challenges: Manual review becomes impractical with increasing volume and complexity of conversations.
  • Solution: Extract structured data from conversations and perform traditional data analysis.
  • Process:
    1. Extract metadata from conversations, such as summaries, tools used, errors, satisfaction, and frustration levels.
    2. Embed the extracted metadata.
    3. Cluster the embeddings to identify segments and themes.
    4. Analyze the clusters to understand user behavior and identify areas for improvement.
  • Cura Library: A tool for summarizing, clustering, and analyzing conversations.
  • Example: Analyzing fake conversations from Gemini to identify topics, frustrations, and errors.
  • Clustering Results: Identifying themes such as data visualization, SEO content requests, and authentication errors.
  • Actionable Insights:
    • Identify missing tools or functionalities.
    • Improve prompts or workflows.
    • Make data-driven decisions about product roadmap.
  • Impact-Weighted Understanding: Prioritizing development efforts based on the frequency and performance of different use cases.
  • Example: If 40% of conversations are about data visualization and the system performs poorly, prioritize building better data visualization tools.
  • Framework for Product Development:
    1. Define evals.
    2. Find clusters.
    3. Compare KPIs across clusters.
    4. Make decisions on what to build, fix, or ignore.
  • Monitoring and Grouping: Tracking performance metrics across different categories of query types over time.

3. Key Arguments and Perspectives

  • Measure to Manage: You can only improve what you measure.
  • Data-Driven Decisions: Product development should be driven by data and analysis, not just intuition or research.
  • Focus on Retrieval: Improving retrieval is crucial because it's a fundamental issue that LLM improvements alone cannot fix.
  • Earn the Right to Tinker with LLMs: Ensure retrieval is effective before focusing on fine-tuning the LLM.
  • Infrastructure Matters: Often, improving the infrastructure and providing the right tools is more effective than simply making the AI better.

4. Notable Quotes

  • "You can really only manage what you measure." - Attributed to Peter Drucker.
  • "The real marker of progress is your ability to have a high quality hypothesis and your ability to test a lot of these hypotheses."

5. Conclusion

The presentation emphasizes the importance of looking at both the inputs (retrieval) and outputs (conversations) of AI systems to drive systematic improvement. Fast evals provide a quick and inexpensive way to evaluate retrieval performance, while analyzing conversations allows for data-driven product development decisions. By focusing on measurement, experimentation, and understanding user behavior, developers can build better AI products.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.