Practical tactics to build reliable AI apps — Dmitry Kuchin, Multinear

AI EngineerAbout 4 min readAug 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Reliable AI applications
  • Software Development Life Cycle (SDLC) for AI
  • Prompt engineering
  • Model non-determinism
  • Data science approach to AI development
  • Reverse engineering metrics
  • Real-world scenarios for evaluation
  • LLM-based evaluation
  • Continuous experimentation and regression testing
  • Benchmarks for AI solutions
  • Explainable AI

Building Reliable AI Applications: Practical Tactics

The Problem with Traditional SDLC for AI

The speaker begins by outlining the standard software development life cycle (SDLC): design, develop, test, and deploy. While this works for traditional software, AI projects, especially those using Large Language Models (LLMs), present unique challenges.

  • Challenge: Achieving consistent reliability (e.g., moving from 50% success rate to near 100%) is difficult due to the non-deterministic nature of AI models.
  • Challenge: Any change to the solution (code, prompts, models, data) can have unexpected impacts.
  • Challenge: Applying data science metrics (groundedness, factuality, bias) alone is insufficient to determine if the solution is working correctly for users.

Reverse Engineering Metrics: Focusing on User Experience

The speaker argues that the key to building reliable AI applications lies in reverse engineering metrics, starting with real-world scenarios and focusing on product experience and business outcomes.

  • Example: A customer support bot at Tweaks. Instead of focusing on factuality, the most important metric is the rate of escalation to human support. A highly factual but unhelpful answer still results in escalation, which is a failure.
  • Key Point: Metrics should be very specific to the end goal and mimic what users want. Universal evaluations don't work.

LLM-Based Evaluation: A Customer Support Bot Case Study

The speaker uses a customer support bot for a bank as a detailed example.

  • Scenario: The bank has FAQ materials, including instructions on how to reset a password.
  • Process:
    1. Use an LLM (e.g., GPT-3.5) to generate user questions based on the FAQ materials.
    2. Define specific criteria that the answer must meet based on the FAQ.
      • Example: The answer must mention that a mobile validation code (SMS) is required. If the user doesn't have a mobile number, they should be directed to support.
    3. Create numerous evaluations that mimic specific user questions and check if the answers meet the defined criteria.
    4. Account for different user personas, as the same question can be asked in different ways depending on the persona.

Multineer: An Open-Source Platform for AI Evaluation

The speaker mentions Multineer, an open-source platform designed to facilitate this evaluation process. He emphasizes that the approach is more important than the platform itself.

  • Example: Multineer allows you to input a question ("How do I reset my password?"), see the model's output, and define specific criteria to measure the answer's correctness.
  • Key Point: The platform helps iterate and generate variations of the same question to ensure consistent and accurate responses.

The Evaluation-Driven Development Process

The speaker outlines a development process where evaluations are built at the beginning, not the end.

  • Process:
    1. Build a first version of the AI solution (POC).
    2. Define the first version of the evaluations/tests.
    3. Run the evaluations and analyze the results.
    4. Examine the details of each evaluation to understand why it failed or succeeded. Average numbers are not informative.
    5. Make changes to the solution (model, logic, prompt, data) based on the evaluation results.
    6. Continuously improve the solution and the evaluations, adding more tests as needed.
    7. Iterate until a satisfactory baseline/benchmark is reached.

The Importance of Regression Testing

The speaker stresses the importance of regression testing.

  • Key Point: Changes that fix one problem can often break something that used to work. Without comprehensive evaluations, these regressions are difficult to catch.

Benchmarking and Optimization

Once a benchmark is established, it allows for confident experimentation and optimization.

  • Examples:
    • Testing different models (e.g., GPT-4 vs. GPT-3.5).
    • Evaluating the impact of Retrieval-Augmented Generation (RAG).
    • Comparing simpler solutions vs. agentic approaches.
    • Simplifying logic for specific parts of the application.

Solution-Specific Evaluation Strategies

The speaker emphasizes that evaluation strategies must be tailored to the specific AI solution.

  • Customer Support Bot: Use LLMs as judges to evaluate the quality of the responses.
  • Text-to-SQL/Graph Database: Create a mock database with a representative schema and mock data to know the expected results for specific queries.
  • Call Center Conversation Classifier: Use simple matching to determine if the conversation is assigned to the correct rubric.
  • Guardrails: Create evaluations to cover questions that should not be answered, questions that should be answered differently, or questions that are not covered in the available materials.

Key Takeaways

  • Evaluate AI applications the way users actually use them.
  • Avoid abstract metrics that don't measure anything important.
  • Use frequent evaluations to enable rapid progress and minimize regressions.
  • Well-defined evaluations lead to "explainable AI" because you understand exactly what the solution does and how it does it.

Conclusion

The speaker advocates for a shift in how AI applications are developed and evaluated. By focusing on user-centric metrics, continuous experimentation, and comprehensive regression testing, developers can build more reliable and explainable AI solutions. The key is to move away from generic data science metrics and embrace a more nuanced, scenario-based approach to evaluation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.