How to run Evals at Scale: Thinking beyond Accuracy or Similarity — Muktesh Mishra, Adobe

AI EngineerAbout 3 min readJul 23, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Evals: Fundamental approach to writing test cases for measuring AI applications.
  • Synthetic Data: Artificially generated data used to validate application outputs.
  • Data Labeling: Categorizing data into different aspects to cover multiple flows and application prospects.
  • Adaptive Evals: Tailoring evaluations based on the specific application type (e.g., RAG, code generation, agents).
  • Trajectory Evaluation: Assessing the paths taken by agents to execute a flow.
  • Measure, Monitor, Analyze, Repeat: A continuous improvement cycle for evaluating and refining AI applications.
  • Eval Development: Defining evaluations based on specific use cases, similar to test-driven development.

Data-Driven Evals

  • Importance of Data: Data is fundamental to writing effective evals.
  • Data Acquisition:
    • Start small with synthetic data to validate application outputs.
    • Continuously improve the data set based on observed system outputs.
    • Label data to cover multiple flows and application prospects.
  • Multiple Data Sets: One data set is often insufficient; use multiple data sets based on different flows and applications.

Evaluation Goals and Objectives

  • Define Goals: Start by defining the goals and objectives of the evaluation.
  • Modular Design: Design the evaluation with modules defined for each component.
  • Data Handling: Optimize data handling with different data sets for different flows.
  • Flow Testing: Test all flows, outputs, and paths within the application.

Adaptive Evaluation Strategies

  • No Universal Eval: Evals depend on the application type.
  • RAG Applications: Evaluate accuracy, similarity, and usefulness.
  • Code Generation: Measure functional correctness and robustness of generated code against the actual code base.
  • Agents:
    • Trajectory Evaluation: Evaluate the paths taken by agents.
    • Multi-Turn Simulation: Evaluate conversations and interactions.
    • Tool Call Correctness: Check the correctness of tool calls and data generation.

Scaling Evals

  • Caching: Cache intermediate results to improve efficiency.
  • Orchestration and Parallelism: Focus on orchestrating and parallelizing evals.
  • Aggregation: Aggregate results for analysis.
  • Frequency: Run evals frequently and continuously improve.
  • Measure, Monitor, Analyze, Repeat: Implement this cycle for continuous improvement.

Evaluation Strategies and Methodologies

  • Use Case Specific: Adapt methodologies based on the specific use case.
  • No Fixed Strategy: There is no one-size-fits-all strategy.
  • Human-in-the-Loop vs. Automation: Balance human involvement with automation based on the desired speed and fidelity.
  • Process Over Tools: Establish a clear process for running evals, as tools cannot automate everything.

Key Takeaways

  • Evals are Crucial: Evals are the most important aspect of AI application development.
  • Eval Development: Define evals based on use cases, focusing on both positive and negative cases.
  • Data Focus: Emphasize the importance of data in the evaluation process.
  • Continuous Improvement: Measure, monitor, analyze, and iterate continuously.
  • Balanced Approach: Balance fidelity and speed in the evaluation process.

Notable Quotes

  • "Evals is the fundamental approach where you are writing sort of test cases to measure your AI applications."
  • "There is no universal eval."
  • "Measure, monitor, analyze and repeat."
  • "Rely on process over tools."

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.