What everyone gets wrong about evals

Lenny's PodcastAbout 3 min readSep 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Evals: Evaluation metrics and processes for Large Language Models (LLMs).
  • Data Analysis: Examining data to inform the creation of relevant evaluation metrics.
  • Stochasticity: The inherent randomness and variability in LLM outputs.
  • Metrics: Quantifiable measures used to assess LLM performance.
  • Feedback Signal: Data-driven insights that guide iterative improvements to LLMs.

Main Topics and Key Points:

The video emphasizes that evaluations (evals) for Large Language Models (LLMs) should not be treated merely as tests. This is a common misconception, as many developers immediately focus on writing tests without first understanding the underlying data.

Important Examples, Case Studies, or Real-World Applications Discussed:

The video contrasts LLM evaluation with software engineering. In traditional software engineering, there are often clear expectations about how a system should function. However, LLMs have a much larger "surface area" and are inherently stochastic, meaning their outputs are variable and unpredictable.

Step-by-Step Processes, Methodologies, or Frameworks Explained:

The video advocates for a data-driven approach to LLM evaluation. The recommended process is to begin with data analysis to understand the characteristics of the LLM's outputs and identify areas that require specific evaluation. This analysis informs the creation of relevant metrics.

Key Arguments or Perspectives Presented, with Their Supporting Evidence:

The central argument is that data analysis should precede test creation in LLM evaluation. This approach is supported by the fact that LLMs are stochastic and have a large surface area, making it difficult to define meaningful tests without first understanding the data.

Notable Quotes or Significant Statements with Proper Attribution:

  • "It's really important that we don't think of evals as just tests."
  • "...you should start with some kind of data analysis to ground what you should even test."
  • "With LMS, it's a lot more surface area. It's very stochastic."

Technical Terms, Concepts, or Specialized Vocabulary with Brief Explanations:

  • Evals: Evaluation metrics and processes for Large Language Models (LLMs).
  • Stochastic: In the context of LLMs, refers to the random and unpredictable nature of their outputs.
  • Surface Area: The range of possible inputs and outputs for an LLM, which is much larger than traditional software systems.
  • Metrics: Quantifiable measures used to assess LLM performance, such as accuracy, fluency, or coherence.
  • Feedback Signal: Data-driven insights that guide iterative improvements to LLMs.

Logical Connections Between Different Sections and Ideas:

The video establishes a clear connection between data analysis, metric creation, and iterative improvement. Data analysis informs the creation of relevant metrics, which in turn provide a feedback signal for improving the LLM.

Any Data, Research Findings, or Statistics Mentioned:

The video does not explicitly mention specific data, research findings, or statistics. However, it implicitly refers to the inherent variability and unpredictability of LLM outputs, which is a well-documented characteristic of these models.

Brief Synthesis/Conclusion of the Main Takeaways:

The main takeaway is that LLM evaluation should be a data-driven process that begins with data analysis to inform the creation of relevant metrics. This approach is essential for effectively evaluating and improving LLMs, given their stochastic nature and large surface area. By using evals to create metrics, developers can confidently improve their applications with a feedback signal to iterate against.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.