Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work — Dat Ngo, Aman Khan, Arize

AI EngineerAbout 5 min readJun 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Observability: Understanding what an AI system is actually doing.
  • Evals: A method for discerning signal (what's going well) from noise (what's not going well) in AI systems.
  • LLM as a Judge: Using a large language model to evaluate the output of another process, including another LLM.
  • Golden Data Set: A high-quality, manually labeled dataset used to evaluate and tune LLM-based evaluations.
  • Code Evals (Heuristics): Using code-based logic or heuristics for evaluations, which is often cheaper and faster than using LLMs or humans.
  • AI Engineering Virtuous Cycles: Iterating on both the AI system itself and the evaluation process to improve quality.
  • Routers: Components in AI architectures that direct traffic to different models or services.
  • Conditional Evals: Evaluating components based on the outcome of previous components, saving resources by avoiding unnecessary evaluations.
  • Agent Evaluation: Evaluating the performance and failure modes of AI agents, often involving complex pathing and tool usage.
  • Trajectory Evaluation: Evaluating the sequence of steps or components an AI agent takes to complete a task.
  • Guardrails: Systems designed to mitigate risks in AI applications, often implemented as inline evaluations.
  • OpenTelemetry (OTEL): A standard for collecting and propagating telemetry data across distributed systems.
  • Log Probability (Log Prop): A measure of confidence provided by some auto-regressive models, useful for evaluating LLM-based evaluations.
  • Meta-Prompting: Using an LLM to automatically optimize prompts based on data, evaluations, and failure analysis.

Observability and Evals: The Foundation of AI Engineering

  • Observability: Answers the question, "What is the thing I built actually doing?" It can manifest as traces, conversations, or analytics.
  • Evals: Are a form of signal, helping to understand what's going well and what's not. They are crucial because manually inspecting every trace is not scalable.
  • LLM Teams: Are often split into platform teams (focused on infrastructure, cost, latency) and application teams (focused on business needs and evals).

The Spectrum of Evals: Beyond LLM as a Judge

  • LLM as a Judge: Involves using an LLM to provide feedback on a process, such as RAG (Retrieval-Augmented Generation). Every step in RAG can be evaluated (e.g., rag relevance).
  • Human Evaluation: User feedback is "incredible signal" for evaluating AI systems.
  • Golden Data Sets: Manually graded datasets that serve as a source of truth. They can be used to tune LLM as a judge by quantifying how well the LLM approximates the trusted data.
  • Code Evals (Heuristics): Cheaper and faster than LLMs or humans. Examples include checking for keywords, regex patterns, or parsable JSON.

The AI Engineering Virtuous Cycles

  • Cycle 1: Building a Better AI System:
    • Collect data (observability, traces).
    • Run evals to discern signal.
    • Identify areas where things went right or wrong (e.g., hallucinations, rag strategy issues).
    • Annotate datasets to verify eval accuracy.
    • Update prompt templates, models, or agent orchestration.
  • Cycle 2: Tuning Evals:
    • Collect failures where the eval was incorrect.
    • Improve the eval prompt template to be more specific and accurate.
  • Velocity: The faster you iterate through these cycles, the better the AI product will be.

Architectures and Eval Complexity

  • Routers: Are components in AI architectures that direct traffic. Evals can be performed on individual components or the entire workflow.
  • Eval Granularity: Evals can be zoomed in to specific LLM calls or zoomed out to evaluate entire agents or workflows.
  • Control Flow: If components have control flow, evaluate the control flow first. Use conditional evals to avoid evaluating downstream components if the control flow is incorrect.
  • Sessions: Evals can be run at the session level to understand overall customer experience.
  • Customization: "Don't use out-of-the-box evals; you'll get out-of-the-box results." Customize evals heavily based on the specific application.

Agent Evaluation: A Deeper Dive

  • Agent Evaluation Focus: Not just "is my AI agent good or bad?" but "what are the failure modes in which my agent fails?"
  • Agent Graph: A framework-agnostic way to visualize and analyze agent pathing across aggregate traces.
  • Tool Usage Analysis: Understanding how often an agent calls specific tools and the evals associated with specific paths.
  • Trajectory Evaluation:
    • Reference Trajectory: The expected sequence of components or steps.
    • Evaluation Methods:
      • Using an LLM to grade the trajectory based on the expected and actual paths.
      • Checking if specific trajectory strings were hit.
      • Matching nodes and their descriptions to the correct trajectory.

Inline Evals (Guardrails)

  • Guardrails: Mitigate risk but come with costs (latency, complexity).
  • System 1 vs. System 2: System 1 is the core orchestration system; System 2 is the guardrail system.
  • Guardrail Limitations: Guardrails are not infallible and act like unit tests for known knowns. Observability and evals are needed to address unknown unknowns.
  • Root Cause Analysis: "A lot of people mistake guardrails as like the thing that needs to be adjusted... Reality is you need to adjust system one."

Asynchronous Processes and OpenTelemetry

  • OpenTelemetry (OTEL): Is crucial for tracing asynchronous processes across multiple services.
  • OTEL Propagation: Allows you to see the flow of data across different applications and services.

Confidence Scores on Evals

  • Log Probability (Log Prop): If using auto-regressive models, log prop can be used as a pseudo-confidence score for eval labels.
  • Encoder-Only Models: Provide a probability of classification.
  • Toolbox Approach: Use a combination of tools to discern where things go well or not well.

Auto-Optimization and Meta-Prompting

  • DSPY: A framework for creating less fragile prompts that work across different models.
  • Meta-Prompting: Using an LLM to automatically optimize prompts based on data, evaluations, and failure analysis.

Conclusion

The presentation emphasizes the importance of a comprehensive approach to AI engineering, focusing on observability, evals, and iterative improvement. It highlights the need to move beyond simple LLM-as-a-judge evaluations and embrace a wider range of techniques, including human feedback, golden datasets, and code-based heuristics. The speaker stresses the importance of customizing evals to the specific application and continuously tuning them based on observed failures. Furthermore, the presentation delves into the complexities of agent evaluation, emphasizing the need to understand failure modes and analyze agent pathing. Finally, the discussion touches on the role of guardrails, the use of OpenTelemetry for tracing asynchronous processes, and the potential for auto-optimization and meta-prompting to further streamline the AI engineering process.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.