Key Concepts
Voice AI, Interruption Problem, Turn Taking, Speech-to-Text (STT), Voice Activity Detection (VAD), Large Language Model (LLM), Text-to-Speech (TTS), Semantic Prediction, Syntax, Prosody, Full Duplex Models, Cascading Model, End-of-Utterance Detection.
Voice AI's Interruption Problem
The primary issue in voice AI is unwanted interruptions. While annoying in general applications like ChatGPT, interruptions in specialized applications (e.g., a voice AI dental assistant) can lead to user abandonment and financial consequences for developers. Turn taking, the system that governs who speaks, is inherently complex.
- Turn Taking Definition: The unspoken system determining who controls the floor in a conversation.
- Complexity: Turn taking is fast and varies across cultures and individuals.
- Cultural Differences: A study showed Danes take longer to respond than Japanese speakers.
Current Handling of Turn Taking in Voice AI
Current voice AI systems use a simplified pipeline for turn taking:
- Speech Input: User speaks.
- Speech-to-Text (STT): Audio is transcribed into text.
- Voice Activity Detection (VAD): Determines if the user has finished speaking.
- LLM Processing: If the user is done, the transcript is sent to an LLM.
- Chat Completion: LLM generates a response.
- Text-to-Speech (TTS): The response is converted to audio.
- Audio Output: The AI's audio response is played to the user.
Voice Activity Detection (VAD) Details:
- Machine Learning Model: A neural network detects speech vs. non-speech.
- Silence Algorithm: If no speech is detected for a set duration (e.g., 0.5 seconds), the user is considered finished.
Learning from Human Conversation
Human turn taking is a "psycho-linguistic puzzle." Humans respond in approximately 200 milliseconds, while speech generation takes around 600 milliseconds, implying prediction is involved.
- Prediction: Listeners predict the end of a speaker's turn.
- Primary Inputs for Prediction:
- Semantics (Content): The most important factor.
- Syntax (Sentence Structure)
- Prosody (Expressiveness, Tone)
- Visual Cues
Three-Stage Model of Human Turn Taking:
- Semantic Prediction: Inferring the speaker's intended message.
- Refinement: Refining the endpoint prediction based on semantics and syntax.
- Finalization: Finalizing the prediction using prosody and acoustic features.
Full Duplex Processing: Humans process input and generate output simultaneously. A figure from a research paper illustrates comprehension and production tracks happening in parallel.
New Approaches for Turn Taking in Voice AI
Current voice AI systems are too simplistic compared to human conversation. New approaches involve augmenting the VAD with models that consider semantics, syntax, and prosody.
1. Cascading Model Augmentation:
- Prevailing Model: STT -> VAD -> LLM -> TTS
- Augmentation: Enhancing the VAD with models that analyze semantics, syntax, or prosody.
Example: LiveKit's Text-Based Semantic Model:
- Input: The last four turns of the conversation (AI, User, AI, User).
- Model: Transformer model.
- Prediction: Predicts the "end of utterance" token.
- Action: If the model indicates the user hasn't finished, the silence algorithm is extended.
Demo: A demo compared a traditional VAD with LiveKit's semantic end-of-utterance model, showing a significant reduction in interruptions.
2. Acoustic Feature Integration:
- Input: Audio tokens.
- Output: Probability of the user having finished speaking.
- Examples:
- Quinn and the Daily team's open-weight smart turn model: Combines transformer and acoustic analysis.
- AssemblyAI's Streaming Speech-to-Text: Emits both transcript and likelihood of speaker finishing.
- QAI's recently released model.
Limitation: Speech-to-text built-in end-of-utterance models only see the user's speech, not the agent's.
3. Full Duplex Models:
- Concept: Models that process input and generate speech simultaneously, similar to the human mind.
- Training: Trained on raw audio data.
- Analogy: Similar to the shift in computer vision from hand-written algorithms to neural networks trained on raw image data.
- Example: Moshi Model: Always listening and generating output, even emitting "natural silence" when not speaking.
- Example: SyncLLM (Meta AI): Forecasts what the user will say about five tokens (200 milliseconds) ahead.
- Limitations: Current full duplex models are optimized for turn taking but are "dumb LLMs" with limited instruction-following capabilities.
Future Predictions
Smarter VAD augmentations and faster models in the cascade pipeline are more likely to solve the interruption problem than full duplex models for commercial use cases. Computers and LLMs think differently than humans, so voice AI may not replicate human mechanisms exactly.
Q&A Highlights
- Visual Cues: Semantics are more important than visual cues in predicting the end of a turn.
- Cost of Voice AI Calls: Depends on conversation length. Calculators are available online. Text-to-speech is the most expensive part of the pipeline.
- LiveKit Model Availability: Available for use via LiveKit's quick start guide.
- Benchmarking: The industry lacks a good benchmark for turn taking.
- Back Channeling: Current approaches use simple VAD logic. Future models could differentiate between back channels and interruptions. Meta AI's full duplex model can natively back channel.
Synthesis/Conclusion
The interruption problem in voice AI is a complex challenge rooted in the difference between simplified AI systems and the nuanced turn-taking mechanisms of human conversation. While current systems rely on basic voice activity detection, advancements are being made through semantic and acoustic augmentations to VADs. Full duplex models offer a promising direction by mimicking the simultaneous processing of the human brain, but practical limitations remain. The future likely lies in refining existing cascade models with smarter VAD augmentations, leveraging faster processing speeds, and acknowledging the unique computational approaches of AI.
AI summaries can miss context or contain errors. Check important details against the original video.





