Key Concepts
- LLM Evaluations (Evals): Systematic measurement of quality in an AI system.
- Agentic AI Applications: AI applications that can perform tasks autonomously.
- Deterministic vs. Nondeterministic: LLMs are nondeterministic, meaning their outputs can vary even with the same input.
- Analyze, Measure, Improve Lifecycle: Iterative process for continuous improvement of AI systems.
- Unit Tests (Level 1 Evals): Fast, automated assertions to check core functionality.
- Human and Model Evaluation (Level 2 Evals): Systematic review and automated critique of quality, involving both human experts and LLMs.
- AB Testing (Level 3 Evals): Real-user experiments to measure the business impact of different system configurations.
- Reference-Based Metrics: Evaluation metrics that compare the output against a known "golden answer."
- Reference-Free Metrics: Evaluation metrics that assess qualities like tone, safety, and helpfulness without a specific correct answer.
- LLM as a Judge: Using a language model to critique system outputs, aligned with human judgment.
- Evaluation-Driven Workflow: A system where improvement is systematic, not accidental.
LLM Evaluations: A Crash Course
Introduction
Building agentic AI applications is challenging, and proper evaluations are crucial for success. Reports suggest high failure rates for AI projects, highlighting the need for reliable development and improvement strategies. This training introduces a strategy for systematically setting up LLM evaluations, aiming to help developers become top-tier AI engineers. The training is derived from a GenAI accelerator program and covers theory, tools, and code examples. Resources, including a GitHub repository, an Excel tool, and PDF slides, are available via a link in the video description.
The Purpose of LLM Evaluations
LLM evaluations help answer critical questions during AI development:
- How does adjusting the system prompt affect performance?
- Is the RAG (Retrieval-Augmented Generation) pipeline working correctly?
- Is the classification step accurate?
- Is the tone of voice appropriate?
- How will the system behave with real-world user inputs?
- Are unsafe or problematic outputs being caught?
- Is system performance degrading in production?
- Is the system robust against prompt injections?
Without evaluations, these questions remain unanswered, leading to potential issues in production. Evaluations help address the nondeterministic and context-sensitive nature of LLMs, where responses can be factually correct but still inappropriate.
Core Challenges in LLM Development
Three core challenges hinder effective LLM development:
- Understanding the Data: Difficulty understanding user inputs, system responses, and edge cases at scale. Testing isolated examples differs significantly from real-world usage.
- Specifications Gap: The gap between desired system behavior and what can be specified in prompts and code. Translating human judgment of good outputs into clear instructions is difficult.
- Inconsistent Behavior: LLMs performing well on test cases but unpredictably on real user inputs. Small changes in wording can lead to drastically different outputs.
Key Components for Success
Success in building AI systems depends on rapid iteration, which requires three key components:
- Evaluation Quality: Systematic measurements of workflow performance through automated tests, human evaluations, and success metrics.
- Debugging Capabilities: Tools and scripts to understand and diagnose failures, including trace logging, data inspections, and error analysis (Langfuse assists with this).
- Change Behavior: Techniques to improve the system based on insights from evaluation and debugging.
Most teams focus only on improving the system (component 3), hindering progress beyond the demo stage. Combining all three creates a virtuous improvement cycle.
What is an Evaluation (Eval)?
An evaluation is a systematic measurement of quality in an AI system. An "eval" is a single metric that measures a specific aspect of performance. Systems can have multiple evals, such as factual accuracy, tone appropriateness, instruction following, and output format compliance. Evals can be used for background monitoring, guardrails, and improvement tools.
The Analyze, Measure, Improve Lifecycle
This iterative cycle addresses the core challenges of data understanding, specification gaps, and inconsistent behavior through systematic evaluation.
- Analyze: Collect examples and categorize failure modes (e.g., user complaints, errors in the system). Langfuse can be used for this.
- Measure: Translate insights into quantitative metrics (e.g., boolean values, rankings, scores).
- Improve: Refine prompts, experiment with different models, and adjust the overall architecture.
Three Levels of Evaluations
Evaluations are categorized into three levels, each with increasing cost and effort:
- Level 1: Unit Tests: Fast, automated assertions run on every code and prompt change.
- Level 2: Human and Model Evaluation: Systematic review and automated critique of quality, run weekly or bi-weekly. Requires human involvement, making it more costly.
- Level 3: AB Testing: Real-user experiments to measure business impact, run with major releases. The most costly level.
Level 1: Unit Tests (Detailed)
Unit tests are fast, automated assertions. In Python, the assert statement is used to create boolean checkmarks. For LLMs, unit tests are typically created around structured output.
Example:
- Input: "My credit card was charged twice this month."
- Workflow Step: Categorize the ticket.
- Output:
{"category": "billing", "confidence": 0.95}
Assertions:
assert result["category"] in ["billing", "technical", "general"]assert isinstance(result["confidence"], float)assert 0 <= result["confidence"] <= 1assert result["category"] == "billing"
Raw events are stored in JSON format within an "evals" folder in the codebase. Python tests are created to assert specific conditions based on these events.
Level 2: Human and Model Evaluation (Detailed)
This level involves both human experts and LLMs. It's crucial to align automated evaluations with human judgment.
Process:
- Human Evaluation First: Humans (ideally domain experts) review data and evaluate system outputs. Tools like Excel sheets or Langfuse can facilitate this.
- LLM as a Judge: Once a good understanding of human evaluation is achieved, an LLM is used to critique system outputs. The most powerful model available should be used (e.g., GPT-4.1).
- Focus on Quality Dimensions: Evaluate specific aspects like factual correctness and helpfulness.
- Generate Detailed Critiques: The LLM should explain why an output is good or bad, not just provide a score.
- Track Human-Model Agreement: Continuously monitor the alignment between human and model evaluations.
Alignment Process:
- Collect model predictions.
- Generate model critiques.
- Get human evaluations on the same data.
- Compare model vs. human judgments.
- Iterate on the evaluator prompt.
- Repeat until sufficient agreement is achieved.
Example Judge Prompt:
"Evaluate this customer service response on these criteria: accuracy, helpfulness, tone. Provide a detailed critique explaining your reasoning. Then rate either 1 or 0 (good or bad)."
Practical Implementation (Excel Example):
- Collect Sample Data: 100 input examples (real or synthetic).
- Run Through System: Obtain model responses for each input.
- Model Critique (LLM): Use a prompt to have the LLM critique each input-output pair and assign a score (good/bad).
- Human Critique: Have a human expert evaluate the same input-output pairs and assign a score.
- Compare and Align: Calculate the agreement score between the model and human evaluations.
- Iterate: If the agreement score is low, refine the LLM's prompt using meta-prompting (asking a powerful model to optimize the prompt based on the discrepancies).
Level 3: AB Testing (Detailed)
AB testing involves running experiments with real users to measure the business impact of different system configurations (e.g., different prompts, models, or workflows).
Implementation:
- Create an AB test on a specific prompt or model.
- Measure the impact on user satisfaction, task completion rate, time to resolution, user engagement metrics, or business outcomes.
- Gather user feedback (e.g., user satisfaction ratings, thumbs up/down).
AB testing is challenging to implement and is most suitable for mature applications with a large user base.
Types of Evaluation Metrics
Two major categories of evaluation metrics:
- Reference-Based Metrics: Compare the output against a known "golden answer." Examples include exact string matching, semantic similarity, code execution results, SQL query correctness, and structured data validation. Easier to implement.
- Reference-Free Metrics: Assess qualities like tone, safety, and helpfulness without a specific correct answer. Examples include tone appropriateness, length constraints, no hallucinations, format compliance, safety, and toxicity. More challenging to implement.
Common Mistakes to Avoid
- Tool-First Thinking: Jumping to new tools without understanding the underlying problem.
- Fix: Start simple and build custom solutions based on specific needs.
- Generic Metrics Obsessions: Drowning in meaningless scores (e.g., helpfulness 4.2, truthfulness 4.5).
- Fix: Focus on specific, actionable metrics. Start with boolean values.
- Avoiding Your Data: Relying solely on tools and ignoring real user data.
- Fix: Look at lots of real data constantly.
- Unaligned LLM Judges: Assuming the LLM judge works without validation.
- Fix: Always validate judge alignment first by bringing a human in the loop.
Key Principles for Success
- Start simple and specific.
- Focus on the biggest problems first.
- Use existing tools before buying new ones.
- Build domain-specific evaluations.
- Start with unit tests and manual review.
- Look at lots of data.
- Sample broadly, then focus on patterns.
- Never stop examining real examples.
- Remove all friction from viewing data.
- Build custom tools for the domain.
- Automate repetitive evaluation tasks.
- Track metrics over time.
- Only use LLMs to scale (generate test cases, create synthetic data, automate critique and labeling), but always validate with humans.
Evaluation-Driven Workflow
Work towards a system where improvement is systematic, not accidental. Success is indicated by:
- Deploying changes confidently.
- Catching failures before users see them.
- Understanding the system's behavior.
- Improvements compounding over time.
Struggling is indicated by:
- Everything breaking something else.
- Being surprised by user complaints.
- Progress feeling like trial and error.
- Inability to measure if changes help.
Level 1 Evals: Creating Scoped Unit Tests for LLMs (Code Example)
This section provides a code example for creating unit tests for LLMs. The example involves classifying customer service tickets and generating responses.
Setup:
- An "evals" folder with an "events" folder containing sample JSON tickets.
- An OpenAI client setup.
- A function to load events.
- A simple workflow for classification and response generation.
Testing:
- Create test scenarios for each event (e.g., billing, feature request, support).
- Load the event.
- Run the process customer message function.
- Create assertions around the category and response.
Example:
def test_billing_categorization():
event = load_event("billing_ticket.json")
result = process_customer_message(event)
assert result["category"] == "billing"
assert len(result["response"]) > 10
The code loops over the test examples, executes the tests, and catches any assertion errors. The example demonstrates a simple setup that can be expanded as the workflow grows in complexity.
Conclusion
LLM evaluations are crucial for building reliable and effective AI applications. By following the principles and methodologies outlined in this training, developers can systematically improve their systems and achieve superior results. The key is to start simple, focus on specific problems, and continuously iterate based on data and human feedback.
AI summaries can miss context or contain errors. Check important details against the original video.