Key Concepts
- Custom AI Evaluations: Data-driven methods to assess AI model performance based on specific product qualities.
- AI Evaluation Platform: A tool (like Stax) for creating, running, and analyzing AI evaluations.
- Evaluation Project: A container for datasets, AI outputs, human ratings, and evaluators within an AI evaluation platform.
- Evaluators: Methods (pre-built or custom) to automatically grade AI outputs based on specific criteria.
- Hidden Gem: A unique, non-touristy travel recommendation (specific to the example use case).
- System Prompt: Instructions given to the AI model to guide its behavior and responses.
- Data-Driven Evaluation: Using quantitative data and metrics to assess AI performance instead of subjective judgment.
Building AI Travel Agent with Stax: A Step-by-Step Guide
1. Creating an Evaluation Project
- The process begins with creating an evaluation project within Stax.
- This project will house all the data and configurations needed for the evaluation.
- Initial exploration can involve manually testing prompts in a playground environment to understand the AI's behavior.
- A system prompt is defined to instruct the AI model (e.g., "be an inspiring travel companion that recommends hidden gems").
- User inputs (e.g., "foodie destinations in Mexico City") are tested against the model and system prompt.
- Human ratings are provided to assess the quality of the AI's output (e.g., "good hidden gem" or "bad tourist trap").
- Stax automatically saves these examples to build an evaluation benchmark.
2. Data Set Preparation
- Alternatively, existing production data can be uploaded as a CSV file.
- The data set should contain a variety of user inputs and the system instruction to be tested.
3. Generating AI Outputs
- Stax allows choosing models from major providers (e.g., Gemini, GPT) or connecting to custom models (fine-tuned or agent orchestrations).
- Outputs are generated for the entire data set at scale.
4. Evaluating AI Outputs
- Individual outputs can be reviewed, and human ratings can be provided.
- Evaluators are used to automatically grade the AI's output.
- Stax provides pre-loaded evaluators for common criteria (e.g., instruction following).
5. Custom Evaluator Definition
- A key feature of Stax is the ability to define custom evaluators tailored to specific product qualities.
- In the example, a custom evaluator is created to score whether an AI output is a "true hidden gem."
- The custom evaluators are then run at scale against the generated AI outputs.
6. Analyzing Results
- Stax displays evaluator scores for each data point, enabling identification of specific failures.
- Aggregated metrics are provided at the top for overall performance assessment.
- These scores can be used for goaling or comparing different models/prompts.
7. Model Comparison Example
- Gemini 2.5 Flash: HiddenGem score of 86, latency of 11.4 seconds.
- GPT-4.1 mini: HiddenGem score of 62, latency of 6.3 seconds.
- This data allows for informed decisions based on the relative importance of speed versus quality.
8. Iteration and Reusability
- The data set and evaluators can be instantly rerun whenever there is a new model, system prompt, or agent orchestration to test.
- This provides hard data to determine which changes improve performance.
Key Arguments and Perspectives
- Subjective vs. Data-Driven AI Evaluation: The video argues for replacing subjective "gut feelings" with objective, data-driven evaluations.
- Importance of Customization: Generic evaluators are insufficient for assessing unique product qualities; custom evaluators are essential.
- Data-Driven Decision Making: Real data enables informed decisions about model selection, prompt engineering, and other AI development choices.
Notable Quotes
- "Building with GenAI can sometimes feel more like an art than proper engineering." - Sara Wiltberger
- "They turn a subjective manual process into a repeatable, data-driven one. And that's why we built Stax..." - Sara Wiltberger
Technical Terms and Concepts
- Latency: The time it takes for an AI model to generate an output.
- Fine-tuned Model: A pre-trained AI model that has been further trained on a specific dataset to improve its performance on a particular task.
- Agent Orchestration: A system that manages and coordinates multiple AI agents to perform a complex task.
Logical Connections
The video logically connects the problem of subjective AI evaluation to the solution of data-driven evaluation using Stax. It demonstrates how to create an evaluation project, prepare data, generate outputs, define custom evaluators, analyze results, and iterate on the AI model or prompt. The model comparison example provides a concrete illustration of how data can inform decision-making.
Synthesis/Conclusion
The main takeaway is that custom AI evaluations, facilitated by platforms like Stax, are crucial for moving beyond subjective assessments and building high-quality AI products. By defining custom evaluators and using data to compare different models and prompts, developers can make informed decisions and continuously improve their AI systems. The process transforms AI development from an "art" to a more rigorous, data-driven engineering discipline.
AI summaries can miss context or contain errors. Check important details against the original video.





