Evals 101 — Doug Guthrie, Braintrust

AI EngineerAbout 5 min readJun 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Evals: Structured tests to measure the quality, reliability, and correctness of AI systems.
  • Flywheel Effect: A feedback loop created by connecting offline and online evals, enhancing development speed and application quality.
  • Tasks: The code or prompt to be evaluated, ranging from simple prompts to complex agentic workflows.
  • Data Sets: Real-world examples used to evaluate task performance.
  • Scores: Logic behind evals, including LLM-as-a-judge and code-based scores.
  • Offline Evals: Pre-production evaluations for identifying and resolving issues.
  • Online Evals: Real-time tracing of applications in production to diagnose performance and reliability.
  • Human-in-the-Loop: Incorporating human review and user feedback to improve AI systems.

1. Introduction to Evals and Brain Trust

  • Doug Guthrie, a solutions engineer at Brain Trust, introduces the company as an end-to-end developer platform for building AI products, emphasizing the importance of evals.
  • Brain Trust's CEO, Anker Goyle, previously developed similar systems, leading to the creation of Brain Trust.
  • Many companies are already using Brain Trust in production for evals and observability of their GenAI applications.

2. Why Use Evals?

  • Evals help answer critical questions about AI system performance, such as whether changes to the underlying model or prompts improve or worsen the application.
  • Evals provide a rigorous process for building with large language models, addressing the challenge of non-deterministic outputs.
  • Anker Goyle views evals as a proactive tool for development, similar to unit tests but used for "playing offense" rather than just "playing defense."
  • Evals create a flywheel effect, cutting development time and enhancing application quality by connecting real-world user interactions with offline development.

3. Core Concepts of Brain Trust

  • Brain Trust's platform includes prompt engineering, evaluability (logs, human review, user feedback), and the flywheel effect.
  • The platform's playground serves as an IDE for LLM outputs, enabling rapid prototyping.
  • Evals are defined as structured tests that check how well an AI system performs, measuring quality, reliability, and correctness.

4. Ingredients of an Eval

  • Task: The code or prompt to be evaluated, requiring an input and an output.
  • Data Set: Real-world examples used to run the task against.
  • Scores: The logic behind the eval, including LLM-as-a-judge (assessing output quality based on criteria) and code-based scores (heuristic or binary).

5. Offline vs. Online Evals

  • Offline Evals: Pre-production, used for iteration and issue resolution, defining tasks and scores.
  • Online Evals: Real-time tracing in production, logging model inputs, outputs, intermediate steps, and tool calls.
  • Online evals diagnose performance and reliability issues, providing metrics like cost, tokens, and duration.

6. Improving Evals: A Matrix

  • A matrix is presented to guide improvement efforts:
    • Good Output, Low Score: Improve evals.
    • Bad Output, High Score: Improve evals or scoring.
  • The advice is to start small, establish a baseline, and iterate from there, rather than trying to create a perfect data set initially.

7. Components Within the Brain Trust Platform

  • Task: A prompt or agentic workflow within the platform, allowing specification of the underlying model, system prompt, and access to tools.
  • Data Set: Test cases with required input, optional expected output, and metadata for filtering.
  • Scores: Code-based scores (TypeScript or Python) and LLM-as-a-judge scores (using an LLM to assess output based on criteria).
  • Auto Evals: A package with out-of-the-box scores (LLM-as-a-judge and code-based) for quick starts.

8. Tips for Scoring

  • Use higher-quality models for scoring, even if the prompt uses a cheaper model.
  • Break scoring into focused areas (e.g., accuracy, formatting, correctness).
  • Test score prompts in the playground before use.
  • Avoid overloading the score or prompt with context; focus on relevant input and output.

9. Playgrounds and Experiments

  • Playgrounds: Used for rapid iteration, pulling in prompts, agents, data sets, and scores to run evaluations.
  • Experiments: Snapshots in time of evals, tracking performance improvements over time.
  • The new "loop" feature utilizes AI to optimize prompts, leveraging evaluation results to improve performance.

10. Running Evals via SDK

  • Brain Trust offers Python and TypeScript SDKs for defining and running evals in code.
  • Users can define prompts, scores, and data sets in their codebase and push them to the Brain Trust platform.
  • Evals can be run as part of CI/CD processes, with a GitHub action example available in the documentation.

11. Moving to Production: Setting Up Logging

  • Instrumenting applications with Brain Trust code to measure quality on live traffic.
  • Logging model inputs, outputs, and intermediate steps to diagnose performance and reliability issues.
  • The flywheel effect is emphasized, allowing production logs to be added back to data sets for offline evals.
  • Initialization of a logger authenticates the application with Brain Trust and points it to a specific project.
  • Wrapping LLM clients and using trace decorators on functions to capture metrics and customize logs.

12. Online Scoring and Custom Views

  • Configuring scores within the platform to run on incoming logs, with a specified sampling rate.
  • Early regression alerts can be created if scores drop below a certain threshold.
  • Custom views allow filtering logs to focus on specific areas of interest for human review.

13. Human-in-the-Loop

  • Critical for the quality and reliability of applications, providing ground truth for evaluations.
  • Two types of human-in-the-loop interactions:
    • Human Review: Using an interface within Brain Trust to parse through logs and add relevant scores.
    • User Feedback: Collecting feedback from users in the application and creating views to power human review.
  • Sarah from Notion emphasized the importance of a dedicated role for human-in-the-loop interactions, such as a product manager mixed with an LLM specialist.

14. Conclusion

  • Brain Trust provides a comprehensive platform for building, evaluating, and monitoring AI applications.
  • The platform's key features include evals, the flywheel effect, offline and online evaluations, and human-in-the-loop interactions.
  • By leveraging these features, developers can create high-quality, reliable AI systems and continuously improve their performance.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.