Building AI Products That Actually Work — Ben Hylak (Raindrop), Sid Bendre (Oleve)

AI EngineerAbout 4 min readJul 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI product iteration
  • Signals (explicit and implicit)
  • Intents
  • Evals (limitations and misconceptions)
  • Trellis framework (Discretization, Prioritization, Recursive Refinement)
  • Workflows (semi-deterministic)
  • Achievable Delta

Building AI Products That Actually Work

Introduction

Ben Hilac, CTO of Raindrop, and Sid, co-founder of Alie, discuss how to build successful AI products through iteration and a structured approach. The presentation emphasizes the importance of understanding user intents, defining signals, and continuously refining AI experiences.

The Current State of AI Products

  • It's an exciting time because focusing on specific use cases and training small models for specific tasks can lead to exceptional results.
  • Examples like Deep Research (from ChatGPT) demonstrate the potential of focused AI products.
  • However, even leading AI providers like OpenAI (with Codex) and Google Cloud, and Grock ship products with issues, highlighting the challenges in building reliable AI applications.
  • Examples of AI failures:
    • Codex generating incorrect or unhelpful tests.
    • Virgin Money's chatbot threatening customers for using the word "virgin."
    • Google Cloud confusing Azure and Roblox credits.
    • Grock providing irrelevant and offensive responses.
    • Grock failing to find Ben's tweets about AI failures.

Why Building AI Products is Hard

  • Communication is hard: It's difficult to communicate desired outcomes to AI models, especially without sufficient context.
  • Undefined behavior: As AI models become more capable, the number of potential edge cases and unexpected behaviors increases.
  • Integration challenges: Integrating AI products with other tools and data formats introduces new complexities.
  • Inability to define scope upfront: The entire scope of a product's behavior cannot be defined in advance, necessitating iteration.

The Role and Misconceptions of Evals

  • Evals are important for iteration but have limitations.
  • Lie #1: Evals tell you how good your product is: Evals only capture known issues and are easily saturated. Goodhart's Law applies.
  • Lie #2: Using LLMs as judges (e.g., "how funny is my joke?") works: The best companies use highly curated datasets and autogradable evals.
  • Lie #3: Eval production data: Moving offline evals online is expensive, inaccurate, and limited to known issues.

The Importance of Signals

  • To build reliable AI apps, you need signals, which are ground truthy indicators of your app's performance.
  • AI issues lack concrete errors, making signals crucial for identifying and addressing problems.
  • The anatomy of an AI issue consists of signals (implicit and explicit) and intents (user goals).

Defining Signals

  • Explicit Signals: Analytics events sent by the app (e.g., thumbs up/down, copy portion of message, preference data).
  • Implicit Signals: Detected behaviors (e.g., refusals, task failure, user frustration).
  • Example: Raindrop uses the number of correct/wrong search results marked by users to improve search quality via RL.
  • Clustering implicit signals can reveal interesting patterns (e.g., user frustration related to searching for tweets).

Exploring and Refining Signals

  • Explore tags and metadata (properties, models, keywords, intents) to understand the context of signals.
  • Intent significantly impacts the nature of an AI issue.
  • Continuously refine and define new issues by analyzing data, talking to users, and identifying unexpected problems.

Constant IV of App Data

  • Maintain a constant intravenous drip of your app's data.
  • Monitor data through Slack notifications or other means.

Trellis Framework (Sid, Alie)

  • Framework for continuously refining AI experiences to scale to millions of users while maintaining reliability and AI magic.
  • Core Axioms:
    • Discretization: Break down the infinite output space into specific buckets of focus.
    • Prioritization: Rank the buckets by their potential impact on the business.
    • Recursive Refinement: Repeat the process within the prioritized buckets.
  • Six Steps:
    1. Initialize Output Space: Launch an MVP agent to collect user data.
    2. Classify Intents: Categorize user data into intents based on usage patterns.
    3. Convert to Workflows: Create semi-deterministic workflows for each intent. A workflow is a predefined set of steps to achieve a certain output.
    4. Prioritize Workflows: Score workflows based on company KPIs.
    5. Analyze Workflows: Understand failure patterns and sub-intents within workflows.
    6. Recurse: Repeat the process from step 2 within each workflow.
  • Prioritization Scoring Mechanisms:
    • Volume Only (Naive): Focus on workflows with the most volume.
    • Volume * Negative Sentiment Score (Recommended): Prioritize workflows with high volume and negative sentiment.
    • Negative Sentiment * Volume * Estimated Achievable Delta * Strategic Relevance (Informed): Incorporate the estimated achievable delta (potential improvement) and strategic relevance.
  • Estimated Achievable Delta: Score the actual achievable delta you can gain from working on that workflow and improving the product.
  • Benefits of Trellis:
    • Structured workflows that are self-attributable, deterministic, and self-bound.
    • Faster and more reliable team movement.
    • Engineered, repeatable, testable, and attributable AI experiences.

Conclusion

Building successful AI products requires a continuous iteration process driven by data, user feedback, and a structured approach. By defining and monitoring signals, understanding user intents, and using frameworks like Trellis, developers can create reliable and engaging AI experiences that scale.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.