Key Concepts
- Evals (Evaluations) as a measure of organizational value
- Data set engineering and reconciliation with reality
- Scorers as specifications for AI applications
- Context-aware prompting beyond system prompts
- Tool definition and output optimization for LLMs
- Model readiness and adaptability to new releases
- Holistic system optimization (data, task, scoring)
- Brain Trust proxy for model provider flexibility
- Loop: Auto-optimization of evals using LLMs
1. Evaluating the Value of Evals
- Main Point: It's crucial to determine if evals are genuinely beneficial to your organization.
- Three Signs of Success:
- Rapid Model Integration: Ability to launch product updates incorporating new models within 24 hours of their release.
- Example: Notion's ability to integrate new models within 24 hours.
- User Feedback Incorporation: A clear path to integrate user complaints into evals.
- Offensive Use of Evals: Using evals to proactively understand solvable use cases before product launch, not just for regression testing.
- Rapid Model Integration: Ability to launch product updates incorporating new models within 24 hours of their release.
- Actionable Insight: If these signs are absent, there's work to be done on your evals.
2. Engineering Great Evals
- Main Point: Great evals require engineering; they don't arise automatically from synthetic data or generic LLM-as-judge scores.
- Data Sets:
- No data set perfectly aligns with reality.
- The best data sets are continuously reconciled with real-world experiences.
- Data set creation should be treated as an engineering problem.
- Scorers:
- Companies should write and modify their own scoring functions.
- Scorers are like a spec or PRD (Product Requirements Document) for your AI application.
- Using generic scorers is like using a spec for someone else's project.
- Actionable Insight: Invest in custom scoring functions tailored to your specific AI application.
3. The Importance of Context in Prompts
- Main Point: Context is crucial in modern prompting, extending beyond just the system prompt.
- Modern Prompt Structure: System prompt + iterative loop (LLM calls, tool calls, incorporation of tool calls).
- Token Distribution: The majority of tokens in a prompt often come from sources other than the system prompt (e.g., tool outputs).
- Tool Definition: Tools should be designed based on what the LLM wants to see, not just as reflections of existing APIs.
- Example: Changing a tool's output from JSON to YAML significantly improved performance due to YAML's token efficiency for LLMs.
- Actionable Insight: Carefully craft tool definitions and outputs to maximize LLM effectiveness.
4. Model Readiness and Adaptability
- Main Point: Be prepared for new models to change everything and engineer your product and team accordingly.
- Example: Brain Trust's internal feature, which was not viable with GPT-4o, became viable with Claude 4 Sonnet due to improved performance on an eval.
- Brain Trust Proxy: A tool (or similar tools) that allows you to switch between model providers without changing code.
- Actionable Insight: Create ambitious evals that can be easily tested with new models as they are released.
5. Holistic System Optimization
- Main Point: Optimize the entire AI system, not just the prompt.
- System Components: Data (for evals), Task (prompt, agentic system, tools), and Scoring Functions.
- Benchmark Result: Optimizing the entire system (prompt, data set, scores) yields significantly better results than optimizing the prompt alone.
- Brain Trust Loop: A new feature that auto-optimizes evals within Brain Trust, allowing users to improve prompts, data sets, and scores.
- Actionable Insight: Consider the entire system when optimizing for performance.
6. Q&A Highlights
- Overfitting Evals:
- The speaker is more concerned about overfitting to a data set without user feedback.
- Human judgment is important when incorporating user feedback into evals.
- Changes with New Models:
- Some use cases (e.g., classifying movie quotes) have worked well since GPT-3.5.
- Other use cases (e.g., complex agentic tasks) may only become viable with newer models.
- The key is to have evals in place to quickly assess the impact of new models on ambitious use cases.
Synthesis/Conclusion
The key takeaways are that effective evals are crucial for AI product development, but they require careful engineering and a holistic approach. This includes continuously reconciling data sets with reality, crafting custom scoring functions, optimizing tool definitions and outputs for LLMs, and being prepared to adapt to new models. The Brain Trust platform offers tools like the Brain Trust proxy and Loop to facilitate these processes, enabling users to rapidly integrate new models and optimize their AI systems for maximum performance. The future of evals involves leveraging LLMs to automate the iterative improvement process, reducing manual labor and accelerating development cycles.
AI summaries can miss context or contain errors. Check important details against the original video.