Key Concepts
- Evals: Evaluations of Large Language Models (LLMs) at the application layer, focusing on user experience and data relevance.
- Vzero: A full-stack visual coding platform by Vercel.
- AI SDK: A software development kit used for building AI applications.
- Prompt Engineering: Designing effective prompts to elicit desired responses from LLMs.
- Hallucinations: Instances where LLMs generate incorrect or nonsensical outputs.
- Data (Eval): The specific input or question used to test the LLM (e.g., "How many Rs are in strawberry?").
- Task (Eval): The method or configuration used to process the input, including system prompts, pre-processing steps, and retrieval-augmented generation (RAG).
- Score (Eval): A metric indicating the success or failure of the LLM's response.
- Deterministic Scoring: Using clear, objective criteria for evaluating LLM outputs (e.g., pass/fail).
- CI (Continuous Integration): Integrating evals into the software development pipeline to automatically assess changes.
- Brain Trust: A platform for evaluating and comparing LLM performance.
- Observability: The ability to monitor and understand the internal state of a system.
- RAG (Retrieval-Augmented Generation): A technique to improve LLM accuracy by grounding it with external knowledge.
The Fruit Letter Counter Story
The talk begins with a story about building a simple app called "Fruit Letter Counter" using Vzero and AI SDK. The initial implementation, using GPT-4, seemed to work perfectly in initial tests. However, after launching the app, users reported incorrect results, highlighting the inherent unreliability of LLMs. This unreliability can render AI applications unusable, even if 95% of the app functions correctly.
Example: The app initially returned the correct number of "R"s in "strawberry" during testing. However, users later reported incorrect counts.
The Basketball Court Analogy for Evals
To visualize and explain the concept of evals, the speaker uses a basketball court analogy:
- Data (Point on the Court): Represents a specific user query or input.
- Task (Shooting the Ball): Represents the method used to process the input (e.g., prompt, model).
- Score (Making the Basket): Indicates whether the LLM's response was correct (blue) or incorrect (red).
- Basket (Glowing Golden Circle): Represents the desired outcome or correct answer.
The distance from the basket represents the difficulty of the query. Points outside the court boundaries represent irrelevant or out-of-scope queries.
Example:
- "How many Rs in strawberry?" - Blue, close to the basket (easy).
- "How many Rs in strawberry banana pineapple mango kiwi dragon fruit apple raspberry?" - Red, farther from the basket (difficult).
- "How many syllables are in carrot?" - Out of bounds (irrelevant).
Building Effective Evals
The speaker emphasizes the importance of understanding your "court" (i.e., the range of inputs your users will provide) when building evals.
Key Steps:
- Understand Your Court: Identify the boundaries of your application's domain and the types of queries users are likely to make.
- Collect Data: Gather real-world user queries to create a representative dataset for evals. Methods include:
- Thumbs up/thumbs down feedback.
- Analyzing logs.
- Monitoring community forums.
- Scouring social media (e.g., Twitter).
- Plot Your Data: Visualize the performance of your LLM across different types of queries. Identify areas where the model struggles (red zones).
- Factor Constants: Separate constant data (e.g., the basic question) from variable tasks (e.g., system prompt, RAG). This improves clarity, reuse, and generalization.
- Share Code: Use middleware (e.g., AI SDK middleware) to share pre-processing logic between evals and production code. This ensures that the practice environment closely mirrors the real-world application.
Traps to Avoid:
- Out-of-Bounds Trap: Spending time on evals for irrelevant queries.
- Concentrated Set of Points: Failing to test across the entire range of possible inputs.
Scoring Evals
Scoring is a crucial step in the eval process. The speaker recommends:
- Deterministic Scoring: Favoring clear, objective criteria (e.g., pass/fail) to simplify debugging and collaboration.
- Simplicity: Keeping scores as simple as possible to ensure they are easily understood and maintained.
- Human Review: Using human review when automated scoring is too difficult.
Trick for Easier Scoring:
- Adding extra prompts to the original prompt to make string matching easier. For example, "Output your final answer in these answer tags."
Integrating Evals into CI
Integrating evals into the CI pipeline allows for automated assessment of changes. Brain Trust provides eval reports that show improvements and regressions across the dataset. This helps developers understand the impact of their changes on the overall performance of the application.
Example: A PR that changes the prompt can be evaluated to see if it improves performance in some areas while degrading it in others.
Conclusion
Evals are essential for building reliable AI applications. By treating evals as practice, developers can systematically improve the performance of their LLMs, leading to better reliability, higher conversion rates, and reduced support costs. Improvement without measurement is limited and imprecise, and evals provide the clarity needed to systematically improve your app.
Key Takeaway: "Improvement without measurement is limited and imprecise."
AI summaries can miss context or contain errors. Check important details against the original video.





