How to talk to statues — Joe Reeve, ElevenLabs

By AI Engineer

Share:

Key Concepts

  • Vibe Coding: A rapid, iterative approach to software development where developers use LLMs and natural language to "glue" together existing APIs and build functional prototypes quickly.
  • 11 Labs: An audio AI foundation model company providing text-to-speech, speech-to-text, AI music generation, sound effects, and a managed AI agents deployment platform.
  • Voice Interaction Patterns: The design of how users communicate with AI agents, including challenges like latency, interruption handling, and multimodal feedback.
  • Multimodal Conversations: Interfaces that combine voice input with visual cues (UI, diagrams, or text) to increase information density.
  • Agentic Workflows: Systems where AI agents use tools (MCPs - Model Context Protocols) and knowledge bases to perform tasks autonomously.

1. The "Statue App" Case Study

Joe, from the 11 Labs growth team, developed a viral application that allows users to take a photo of a statue and engage in a real-time voice conversation with it.

  • Technical Stack: The app uses OpenAI for deep research on the statue's identity, 11 Labs’ "Voice Design" API to generate a historically appropriate voice, and the 11 Labs agents platform to manage the conversation.
  • Development Speed: The prototype was built in approximately two hours using Cursor (an AI-powered code editor).
  • Impact: The project gained 1.5 million impressions within two days, attracting interest from major museums (e.g., British Museum, Science Museum) and auction houses (e.g., Christie’s, Bonhams).
  • Key Insight: The "hard" part of the project was not the coding, but the "glue" work—connecting APIs and crafting a compelling narrative.

2. Vibe Coding and Rapid Prototyping

  • Methodology: Vibe coding emphasizes experimentation over traditional, rigid software engineering. It allows non-engineers or developers to build complex apps by describing desired outcomes to an LLM.
  • Frameworks: The speaker highlights the importance of using existing, scalable APIs (like those from 11 Labs or Supabase) rather than building infrastructure from scratch.
  • Future Potential: The speaker compares the current state of vibe coding to the early days of Facebook Instant Games, suggesting that a "TikTok moment" for vibe coding—where social sharing and easy creation collide—is imminent.

3. Voice as an Interface: Challenges and Solutions

The discussion identified several friction points in current voice-AI interactions:

  • Information Density: Voice is often too slow for complex data. Users prefer "parallel input/output"—speaking raw intent while receiving high-density visual information (diagrams/text) in return.
  • The Interruption Problem: Humans are often too polite to interrupt AI agents. Joe suggests that "skim listening" (similar to podcast speed-dialing or skipping) and visual cues (e.g., an icon indicating the agent wants to speak) could improve the experience.
  • Philosophical Design: For inanimate objects, the "voice" must be curated. For example, a statue in a British museum that originated in China but was carved in Vietnam requires a nuanced, historically informed voice profile.

4. Content Creation and Virality

Joe shared his strategy for creating viral content:

  • The 80/20 Rule: High-quality editing is accessible via mobile tools like CapCut. He emphasizes that the "hook" must occur within the first 6–12 seconds to prevent user drop-off.
  • Audio-First Strategy: Music is often underrated. Joe recommends either matching speech to a pre-selected vibe or generating music to fit the narrative.
  • Equipment: He uses a high-quality Bluetooth lapel mic (DJI) to ensure professional audio, which significantly improves perceived quality.

5. Synthesis and Takeaways

  • Actionable Insight: The barrier to entry for building sophisticated AI applications has collapsed. Developers should focus on "glue" logic and storytelling rather than solving low-level technical problems.
  • Future Direction: The industry is moving toward multimodal agents that can signal their intent to speak, allow for non-linear navigation of audio (skimming), and provide visual context to supplement voice.
  • Notable Quote: "The hard bit is actually not the user management... it's mostly the relying on our APIs and our agents platform to do the heavy lifting." — Joe, 11 Labs.

Conclusion: The intersection of "vibe coding" and voice AI is creating new cultural experiences, particularly in public spaces like museums. The next phase of development will focus on making these interactions more natural, information-dense, and socially integrated.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video