How to create custom AI evaluators in Stax

Google for DevelopersAbout 3 min readAug 29, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Evaluator Gallery: A Stax feature for managing and creating evaluators.
  • LLM as a Judge/Auto Raters: Using LLMs to automatically evaluate AI outputs.
  • Evaluator Prompt: The instructions given to the LLM judge model.
  • Rating Categories: Predefined categories used by the LLM to grade outputs (e.g., pass/fail, hidden gem/tourist trap).
  • Metric Score and Color Mapping: Assigning numerical scores and colors to rating categories for visualization and aggregation.

Evaluator Gallery and Preloaded Evaluators

Stax provides an "evaluator gallery" for managing and creating evaluators. The platform comes with preloaded evaluators for common Gen AI criteria such as "instruction following" and "verbosity."

The Need for Custom Evaluators

AI developers often require specific and nuanced criteria for evaluating AI outputs tailored to their products and use cases. The example given is an AI travel agent that should recommend unique, "hidden gem" locations.

The Problem with Manual Evaluation

Manually evaluating AI outputs (e.g., "eyeballing" if a travel recommendation feels like a hidden gem) is time-consuming and doesn't scale.

LLM as a Judge (Auto Raters) Solution

Stax offers a solution by allowing users to create their own LLM-based evaluators, also known as "LLM as a judge" or "auto raters," to automate the evaluation process.

Creating a New Evaluator: Step-by-Step

  1. Define the Base LLM: Choose the LLM that will act as the judge model.
  2. Craft the Evaluator Prompt: Write a prompt that instructs the judge LLM on how to score the AI's output. The quality of the prompt is crucial for effective evaluation.
  3. Specify Evaluation Criteria: Be as specific as possible when defining the criteria the LLM should be judging. For example, instead of asking "is this a hidden gem?", describe the characteristics of a hidden gem in detail (e.g., "a non-obvious option, authentic local experience, a specific place, neighborhood, or activity").
  4. Provide Examples: Adding examples to the prompt helps the LLM understand the desired criteria.
  5. Define Rating Categories: Give the evaluator clear rating categories to use when grading (e.g., "hidden gem," "popular favorite," "tourist trap").
  6. Map Rating Categories to Metrics: Assign a metric score and color to each rating category. This mapping defines how results will be visualized and aggregated into metrics within projects and analytics.

Using the Evaluator

Once the evaluator is created, it can be used in projects to score AI outputs.

Validating and Refining the Evaluator

It's important to compare the auto rater's scores against human ratings on a small set of outputs to ensure alignment. If the ratings don't align, the evaluator prompt should be tweaked until the evaluator matches the desired quality bar. This ensures confidence in scaling evaluations.

Conclusion

Stax's evaluator gallery and auto-rater functionality provide a powerful way to automate and scale the evaluation of AI outputs. By crafting detailed evaluator prompts and validating the results against human ratings, developers can ensure that their AI models meet specific quality standards and criteria.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.