The Future of Evals - Ankur Goyal, Braintrust

AI EngineerAbout 3 min readAug 10, 2025Watch original
THE SUMMARYAI-generated

Brain Trust: Evals, Loop, and the Future of AI Evaluation

Key Concepts:

  • Evals: Evaluations of AI model performance, prompt effectiveness, data set quality, and scorer accuracy.
  • Loop: An agent within Brain Trust that automatically optimizes prompts, data sets, and scorers using frontier models.
  • Frontier Models: Cutting-edge large language models (LLMs) like Claude 4.
  • Scorers: Functions or models that assess the quality of AI outputs.
  • Data Sets: Collections of data used to train and evaluate AI models.

1. The Current State of Evals:

  • Brain Trust has observed significant EVEL usage, with the average organization running 13 EVELs daily and some running over 3,000.
  • Advanced companies spend over two hours daily working through their evals.
  • Despite the sophistication of AI products, the evaluation process remains largely manual, relying on dashboards for insights.
  • The primary challenge is translating dashboard insights into actionable changes to code or prompts.

2. Introducing Loop: Automated Optimization:

  • Loop is an agent built into Brain Trust designed to automate the optimization of prompts, data sets, and scorers.
  • Loop's development is based on two years of quarterly evals on frontier models, assessing their ability to improve prompts, data sets, and scorers.
  • Claude 4 is highlighted as a breakthrough model, performing six times better than its predecessors in these optimization tasks.
  • Loop is accessible to existing Brain Trust users via a feature flag and supports various models, including OpenAI, Gemini, and custom LLMs.

3. Loop's Functionality and UI:

  • Loop operates directly within the Brain Trust UI, allowing users to view data and prompts alongside suggested edits.
  • The UI provides side-by-side comparisons of original and optimized prompts, data, and scoring ideas.
  • Users can choose between manual review of suggestions and fully automated optimization using a toggle.

4. Revolutionizing Evals:

  • The speaker expresses excitement about the potential of frontier models to revolutionize evals.
  • The goal is to move from manual evaluation to automated optimization, leveraging the capabilities of advanced AI models.

5. Call to Action and Hiring:

  • Users are encouraged to try Brain Trust and Loop, providing feedback to guide further development.
  • Brain Trust is actively hiring for UI, AI, and infrastructure roles.
  • A QR code is provided for interested individuals to contact the company.

6. Notable Quotes:

  • "The average org that signs up for Brain Trust runs almost 13 EVELs a day."
  • "Claude 4 in particular was a real breakthrough moment...it performs almost six times better than the the previous leading model before it."
  • "EVELs have been a critical part of building some of the best AI products in the world but the task of actually doing evaluation has been incredibly manual."

7. Logical Connections:

  • The presentation begins by highlighting the importance and current limitations of manual evals.
  • It then introduces Loop as a solution to automate and improve the evaluation process.
  • The speaker emphasizes the role of frontier models, particularly Claude 4, in enabling this automation.
  • The presentation concludes with a call to action, encouraging users to try Loop and contribute to its development.

8. Synthesis/Conclusion:

The presentation outlines Brain Trust's vision for the future of AI evaluation, moving from manual processes to automated optimization powered by frontier models. Loop is presented as a key tool in this transformation, offering users the ability to automatically improve prompts, data sets, and scorers. The emphasis on user feedback and ongoing development suggests a commitment to refining Loop and adapting to the rapidly evolving landscape of AI.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.