How to build an accuracy pipeline for your AI app
By Google Cloud Tech
Handling Hallucinations & Achieving Accuracy in LLM Applications
Key Concepts:
- Hallucination: An LLM generating content that is nonsensical or inconsistent with its training data or provided context.
- Inaccuracy: A broader term encompassing any deviation from desired correctness in LLM outputs.
- Synthetic Data: Data generated by an LLM for testing purposes.
- Accuracy Pipeline: A system for testing and evaluating the accuracy of LLM applications.
- Offline Evaluation: Testing conducted outside of real-time user interaction.
- Online Evaluation: Testing conducted during live user interaction.
- Retrieval Augmented Generation (RAG): A technique to improve LLM accuracy by grounding responses in external knowledge sources.
Understanding Accuracy & the Problem Space
The core issue discussed is that applications integrating Large Language Models (LLMs) are prone to inaccuracies, including hallucinations. While seemingly obvious, addressing this requires a systematic approach. The speakers emphasize that defining “accurate enough” is crucial, and this definition should be based on established guidelines or relevant data. This data can take several forms: example prompts with expected answers, data from existing systems (like customer service interactions – with permission), or, surprisingly, data generated by LLMs themselves – referred to as “synthetic data.” The speakers note that generating test data with LLMs is significantly more efficient than manual creation, a practice they’ve been doing for decades. As stated by one speaker, “You can use an LLM to evaluate your AI application. It is just a matter of writing the appropriate prompts.”
Building an Accuracy Pipeline: Testing Methodology
The discussion pivots to a testing framework for evaluating LLM application accuracy. This is framed as an extension of standard software testing practices, specifically Behavior-Driven Development (BDD). The proposed system consists of inputs, an application under test, and outputs. The key is establishing an “accuracy pipeline” to evaluate each stage.
The process involves:
- Inputs: Providing the LLM application with prompts or data.
- Application Under Test: The LLM-powered application processes the input.
- Outputs: The application generates a response.
- Accuracy Evaluation: Assessing the output’s correctness.
Evaluating Accuracy: Methods & Techniques
Evaluating the accuracy of LLM outputs can be achieved through several methods:
- Human Review: Manually assessing responses for correctness and quality.
- Text Similarity Algorithms: Comparing outputs to known “good” outputs using algorithms to measure similarity.
- LLM-Based Evaluation: Utilizing another LLM to analyze the output based on predefined criteria. This is presented as the most effective technique.
The LLM-based evaluation can be structured using a “grading rubric” with categories like:
- Prompt Alignment: Does the response address the user’s initial prompt?
- Grounding in Truth: Is the response based on reliable sources (RAG, provided context)?
- Grammatical Correctness & Tone: Is the response well-written and appropriate?
- Guideline Adherence: Does the response comply with system instructions and guidelines?
This evaluation can be broken down into smaller components for granular analysis. The speakers emphasize that prompting the evaluation LLM effectively is key, framing these prompts as “unit tests” for the AI application.
Offline vs. Online Evaluation
The discussion differentiates between two types of evaluation:
- Offline Evaluation: Testing conducted independently of user interaction, allowing for experimentation with different models or versions. This is useful for comparing performance and cost-effectiveness.
- Online Evaluation: Verifying application quality in a production environment while users are actively using it. This involves monitoring and logging to detect inaccuracies in real-time. This is likened to “assert statements in production code.”
Online Evaluation in Multi-Agent Systems
The speakers illustrate online evaluation within a multi-agent system, consisting of an agent, an orchestrator, and user input/output. Online evaluation can be inserted at multiple points:
- Output Stage: Verifying the final output for accuracy, usefulness, and consistency.
- Agent-Level Evaluation: Evaluating the output of each agent before it’s passed on, ensuring accuracy at every step.
- Orchestrator Plan Evaluation: Assessing the orchestrator’s plan before execution, allowing for adjustments if the plan is flawed.
This layered approach aims to prevent errors from propagating through the system. Early detection allows for retries or requests for user clarification, improving overall reliability.
The Trade-offs & Benefits of Continuous Evaluation
Implementing continuous online evaluation involves a trade-off between increased accuracy and potential performance overhead. While not always beneficial in every part of the system, the speakers advocate for leveraging LLMs to evaluate against a rubric at every stage. Catching errors early prevents them from compounding and impacting the final output. As one speaker states, “when we do evaluate every step of the way, we know that an error that say came from this agent doesn't propagate through to everything else and eventually make it into our output.”
Conclusion
The video concludes by reiterating the core takeaway: LLMs can be effectively used to evaluate the accuracy of other LLM applications. This capability enables scalable and continuous testing, leading to more reliable and trustworthy AI-powered systems. The speakers encourage viewers to explore the resources linked in the description to learn more and implement these techniques. The overall message is one of empowerment – developers now have powerful tools to gain confidence in the accuracy of their LLM-enabled applications.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Building Great Agent Skills: The Missing Manual
AI Engineer

Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov
AI Engineer

Full Claude Guide: Beginner to Pro in Under 15 Minutes
Dan Martell

This Skill Turns Your Agents Into Neckbeards...
NeuralNine

PI AGENT FULL COURSE: Master Pi Agent in 30 Minutes
David Ondrej

Build long-running agents with Google’s Agentic Stack | The Agent Factory
Google Cloud Tech

How to Make Your AI Agent Crash Proof in 1 Install (Free)
corbin