Shipping AI That Works: An Evaluation Framework for PMs – Aman Khan, Arize

By AI Engineer

Share:

Key Concepts

  • The Critical Need for AI Evaluation (Eval): LLMs hallucinate and are unreliable, necessitating rigorous evaluation frameworks beyond traditional software testing.
  • LLM as a Judge, with Human Oversight: Utilizing LLMs for scalable evaluation, but requiring human validation and iterative refinement.
  • Iterative Evaluation & “Evals for Evals”: Continuous improvement of prompts, evaluation metrics, and the evaluation system itself through data analysis and human feedback.
  • Evolving Role of the AI Product Manager: Demanding deeper technical understanding, prompt ownership, and a shift towards data-driven (“thrive coding”) development.
  • Importance of Explainability & Traceability: Understanding why an LLM makes a decision is crucial for trust and debugging, leveraging tools like OpenTelemetry.

The Rise of AI Evaluation & the AI Product Manager

The increasing complexity of AI, particularly Large Language Models (LLMs), demands a new approach to product development. Leading LLM developers like Kevin (OpenAI CPO) and Mike (Anthropic CPO) acknowledge that LLMs hallucinate and are inherently unreliable, representing approximately 95% of the LLM market share. This necessitates a shift from traditional product management to a more technical, data-driven approach, giving rise to the “AI Product Manager” (AIPM) role. The bar has been raised in terms of what’s expected to be delivered, moving beyond quick prototyping (“vibe coding”) to a focus on reliable evaluation (“thrive coding”).

Building an Evaluation Framework

A robust evaluation (eval) framework is central to building trustworthy AI systems. Unlike deterministic software (1+1=2), LLMs are non-deterministic and can be manipulated. Aman outlines a four-component framework for building evals: Role, Context, Goal, and Terminology/Label. The most scalable approach for production environments involves using LLMs themselves as judges, but this requires careful consideration. Initial experiments demonstrate that LLMs are not inherently reliable judges; a robotic response was initially misclassified as “friendly” when evaluated for tone.

Iterative Refinement & Human-in-the-Loop

Relying solely on LLMs for evaluation is insufficient – a “trust but verify” approach is essential. The process involves an iterative loop: developing a prompt and dataset, running an LLM evaluation, manually reviewing and relabeling data, comparing human labels to LLM scores, and refining the LLM evaluation prompt or model. Adding a prompt instruction to offer a discount based on email collection, for example, resulted in the LLM judge accurately identifying the presence of a discount offer in 100% of examples, demonstrating the impact of prompt engineering.

Crucially, understanding why an LLM assigns a particular score is vital. AISE generates explanations for its evaluations, detailing the reasoning behind assessments. Furthermore, the concept of “evals for evals” is introduced – evaluating the evaluation system itself using a code evaluator to compare human labels against LLM judge scores.

Technical Infrastructure & Data Augmentation

AISE leverages OpenTelemetry for tracing and logging, augmenting existing data with metadata (user ID, session ID, etc.) to provide richer context for evaluation. This allows for detailed trace analysis, visualizing the input, output, and metadata of AI requests. Auto instrumentation simplifies the process of adding tracing code to existing applications. The platform supports the implementation of code evaluators (e.g., Python functions) to programmatically compare human and LLM labels.

The Future of AI Product Management

The speaker proposes a shift towards evaluation-based requirements, suggesting that “evals are the new requirements spec.” This necessitates a more technical skillset for Product Managers, including comfort with code and data analysis. The development of AI systems is analogous to building self-driving cars – an iterative process of adding complexity and requiring new data and refinement at each stage. A case study highlights agent-human collaboration, where an agent leverages a human (CFO) as a “tool” when it lacks information. AISE is used by companies like Uber, Reddit, and Instacart, and has received investment from DataDog and Microsoft, positioning it as a leading platform in the AI evaluation space.

Conclusion

The development and deployment of reliable AI systems hinge on robust evaluation frameworks. While LLMs can be leveraged as scalable judges, human oversight, iterative refinement, and a deep understanding of the underlying data are crucial. The role of the AI Product Manager is evolving, demanding a more technical skillset and a shift towards data-driven development. By embracing a “trust but verify” approach and prioritizing explainability, we can move beyond “vibe coding” and build AI systems that are not only powerful but also trustworthy and reliable.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video