The truth about Claude Code's evals

Lenny's PodcastAbout 4 min readSep 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Evals: Evaluations of large language models (LLMs), particularly in the context of code generation and understanding.
  • Claude: A specific large language model (LLM), likely referring to Anthropic's Claude.
  • Error Analysis: The process of identifying, categorizing, and understanding the types of errors an LLM makes.
  • Dogfooding: Internal use of a product or service by its developers to identify bugs and improve its quality.
  • Coding Benchmarks: Standardized tests used to measure the performance of LLMs on coding tasks.
  • Fine-tuned Cloud Models: LLMs that have been specifically trained for cloud-based applications.

Main Argument: The Prevalence and Importance of Evals, Even When Denied

The core argument is that despite strong opinions against formal evaluations ("evals") of large language models (LLMs) on platforms like X (formerly Twitter), evals are implicitly and explicitly crucial for the development and improvement of these models, particularly in the context of code generation.

Evidence and Supporting Details:

  1. Contradictory Stance on Evals: The video highlights the apparent contradiction where some entities, like "Claude code," publicly state they don't do evals, while their products are actually built upon extensive evaluation processes.
  2. Evals Underpin Fine-Tuned Models: The speaker asserts that fine-tuned cloud models are developed based on coding benchmarks. These benchmarks are a form of evaluation.
  3. Systematic Error Analysis: The speaker suggests that companies developing LLMs like Claude likely engage in systematic error analysis. This involves identifying patterns in the model's mistakes to improve its performance.
  4. Monitoring Usage Patterns: The speaker posits that LLM providers monitor user behavior, including the number of users, chat creation rates, and chat durations. This data collection is a form of evaluation, providing insights into how the model is being used and where it might be failing.
  5. Internal Dogfooding as Evaluation: The video emphasizes the importance of internal "dogfooding," where developers use their own LLMs. When issues arise during internal use, they are addressed, which constitutes an implicit form of error analysis and evaluation.
  6. Implicit Error Analysis: The speaker argues that even when formal evals are not conducted, developers are constantly performing error analysis when they fix bugs or improve the model based on user feedback or internal testing.

Examples and Real-World Applications:

  • Claude Code: The example of "Claude code" is used to illustrate the discrepancy between publicly stated policies against evals and the likely reality of ongoing evaluation processes.
  • Coding Benchmarks: The mention of coding benchmarks highlights a specific type of evaluation used to assess the performance of LLMs on coding tasks.

Logical Connections:

The video establishes a logical connection between the development of high-performing LLMs and the necessity of evaluation. It argues that even if formal evals are avoided, various forms of implicit and explicit evaluation are essential for identifying and addressing errors, improving model performance, and ensuring user satisfaction.

Notable Quotes:

  • "X or Twitter is like a medium where you just get all these strong opinions of like don't do evals. It's bad. We tried it. It doesn't work." - This quote highlights the negative sentiment towards evals in some online communities.
  • "We're claud code and we don't do evals." - This quote exemplifies the contradictory stance of some LLM developers.
  • "A lot of these applications are standing on the shoulders of evals." - This statement emphasizes the fundamental role of evals in the development of LLMs.

Technical Terms and Concepts:

  • Evals: Evaluations, assessments, or tests used to measure the performance of LLMs.
  • Claude: A specific LLM, likely referring to Anthropic's Claude.
  • Error Analysis: The process of identifying, categorizing, and understanding the types of errors an LLM makes.
  • Dogfooding: Internal use of a product or service by its developers to identify bugs and improve its quality.
  • Coding Benchmarks: Standardized tests used to measure the performance of LLMs on coding tasks.
  • Fine-tuned Cloud Models: LLMs that have been specifically trained for cloud-based applications.

Synthesis/Conclusion:

The video concludes that despite some public opposition to formal evaluations, evals are implicitly and explicitly crucial for the development and improvement of LLMs, particularly in the context of code generation. Various forms of evaluation, including coding benchmarks, error analysis, user monitoring, and internal dogfooding, are essential for identifying and addressing errors, improving model performance, and ensuring user satisfaction. The claim that some entities "don't do evals" is likely misleading, as these entities are likely engaging in various forms of evaluation, even if they are not explicitly labeled as such.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.