Conversation with Bibo Xu: How agent conversations are evolving with Google AI
By Google for Developers
Key Concepts
- Large Language Models (LLMs): AI models trained on vast amounts of text data, enabling advanced natural language understanding and generation.
- Human-Computer Interaction (HCI): The study of how humans interact with computers, focusing on making these interactions more intuitive and effective.
- Multimodality: The ability of AI systems to process and integrate information from multiple sources, such as text, audio, and vision.
- Far-field Speech Recognition: The ability of a device to accurately understand spoken commands from a distance, without requiring the user to be close to a microphone.
- Echo Cancellation: Technology that allows a device to filter out its own audio output, enabling it to hear the user's voice clearly.
- Word Error Rate (WER): A metric used to measure the accuracy of speech recognition systems, indicating the percentage of words that are incorrectly transcribed.
- Code Mixing: The practice of blending words and phrases from two or more languages within a single conversation or sentence.
- Situated Assistant: An AI assistant that is aware of its physical environment and can interact with users based on what it sees and hears.
- Universal Assistant: An AI agent designed to help users with a wide range of tasks, integrating various capabilities.
- Latency: The delay between an input and its corresponding output in a system. Low latency is crucial for natural and responsive interactions.
- Interruption (Barge-in): The ability of a user to interrupt an AI system while it is speaking, allowing for more natural conversational flow.
- Full Duplex Dialogue: A communication system where both parties can speak and listen simultaneously, mimicking human conversation.
- AGI (Artificial General Intelligence): A hypothetical type of AI that possesses the ability to understand, learn, and apply knowledge across a wide range of tasks at a human level.
Evolution of Voice Systems and Human-Computer Interaction
The discussion centers on the significant evolution of voice-based AI systems, moving from basic command-and-control functionalities to sophisticated, fluid dialogue capabilities. This evolution is largely driven by advancements in Large Language Models (LLMs).
Early Days of Voice Interaction
- Initial Focus: Early voice systems, like those at Bell Labs, focused on simple tasks such as recognizing spoken numbers.
- Google Home and Far-field Speech Recognition: The development of Google Home marked a significant step with the implementation of far-field speech recognition and echo cancellation. This allowed devices to understand users from a distance, a crucial development for in-home assistants.
- Challenges with Text-Based Input: Even with improved speech recognition, user adoption for voice input in traditional text-based interfaces (like search boxes) was challenging. The mental model for users was not yet established for this modality.
- Word Error Rate (WER) Improvement: A key technical advancement was the reduction in WER. When the speaker started, WER was around 20%, making systems unusable. By 2015-2016, it had dropped to 5-8%, making systems "completely usable."
The Shift Towards Fluid Dialogue with LLMs
- Brittle Command and Control: Previous voice systems were often "brittle," relying on form-filling-like interactions. Deviations from expected phrasing would cause the system to break.
- LLMs as the Catalyst: LLMs are identified as the core technology enabling "human-level dialogue systems." They allow for free-flowing conversations where the AI understands context and nuances.
- User Experience and Environment: The comfort level of users with voice interaction is influenced by the environment. While public use remains challenging, home and car environments have proven more conducive, especially for quick commands where pulling out a phone is inconvenient.
- Beyond Commands: True Dialogue: The current paradigm shift is from "brittle systems" to "free open dialogue," exemplified by Gemini Live. This allows for more natural, human-like conversations where users can discuss problems and receive advice.
Multimodal AI and Project Astra
- Gemini Multimodeling: Bibbo Shu leads product management for Gemini multimodal modeling, focusing on audio and vision capabilities. This includes native audio dialogue, real-time video understanding, expressive speech generation, and translation.
- Project Astra: This research prototype explores the capabilities of a "universal assistant" built on Gemini. The goal is to create an AI agent that can assist users with a wide range of tasks.
- Situated Assistant Concept: Astra is described as a "situated assistant" that "sees the world with you." It can help users based on visual input, such as assisting with cooking or identifying objects.
- Real-World Application Example: A compelling application demonstrated was helping a blind user navigate the world, showcasing the potential of multimodal AI for accessibility.
Technical Advancements and Challenges
- Naturalness and Quality of Dialogue: LLMs are improving the naturalness of dialogue, making AI sound less robotic and more like a smart conversational partner. This includes incorporating tone, pitch, and inflection.
- Expressive Speech Generation: AI models can now generate speech with tone and pitch, making communication more nuanced and human-like.
- Real-time Audio Translation: A significant advancement is live, end-to-end audio translation. Instead of transcribing to text and then translating, LLMs can directly translate spoken language in real-time, preserving the original speaker's voice characteristics. An example is live speech-to-speech translation in Google Meet.
- Code Mixing and Language Flexibility: LLM-based audio models can handle "code mixing" (blending languages) seamlessly. The AI can understand and respond in different languages within the same conversation, adapting to the user's input.
- Memory and Its Challenges: Integrating capabilities like memory into AI systems presents significant challenges. An example was a bug where Astra "remembered" it couldn't see from a past conversation, even after the visual system was restored. This highlights the complexity of debugging and understanding AI memory.
- Integration of Capabilities: Combining multiple AI capabilities (audio, vision, memory, proactivity) into a single system for a universal assistant is a major research challenge. This integration can lead to unexpected issues and bugs.
- Latency as a Critical Factor: Latency is paramount for natural dialogue. The goal is to achieve sub-500-millisecond latency for human-level conversation.
- Interruption (Barge-in) and Full Duplex Dialogue: Innovations are focused on enabling seamless interruption (barge-in) and full duplex dialogue, where the AI continues to listen even while speaking. This requires sophisticated models that can differentiate between user input and background noise.
- Open Mic Experience: The shift from traditional "mic on/off" systems to an "open mic" where the AI is constantly listening is enabled by models that can distinguish intended user input from background chatter.
Ethical Considerations and User Perception
- Distinguishing AI from Humans: As AI becomes more human-like, there's a growing concern about distinguishing between human and AI interactions. The goal is not to deceive users but to create effective communication tools.
- Emotional Understanding: While AI can pick up on words and expletives, understanding human emotions like anger or frustration is a complex area. The aim is for AI to respond appropriately to user emotions, offering help rather than disengaging.
- User Expectations: Users are increasingly demanding more from AI assistants, expecting them to remember past interactions, understand preferences, and perform actions.
- Politeness and Demands: While not necessarily seeing increased politeness, researchers observe users becoming more "demanding" as they expect more from AI's capabilities, including memory and personalized responses.
- Learning from Mistakes: A key area for future development is for AI to learn from user frustration or anger, understanding what went wrong and how to improve its responses.
Future Directions and Developer Insights
- LLM-Powered Products: LLMs are seen as the key to unlocking a wide range of products and services that leverage fluid communication.
- Developer Opportunities: The live API offers significant potential for developers to build voice agents for companies, integrate vision capabilities for troubleshooting in manufacturing, and create new user experiences.
- Screen Sharing and Digital World Interaction: Vision capabilities extend to the digital world through screen sharing, allowing AI to assist with tasks on a user's computer or phone.
- Execution and Action: A major focus for the future is the "execution part" of being an assistant – enabling AI to actively "do stuff" for users, not just talk about it. This involves controlling computers, using APIs, and automating tasks.
- Transforming Communication: Advancements in AI are fundamentally changing how people communicate, particularly in writing, by enabling users to express ideas more freely and then refine them with AI assistance.
Detailed Summary
The Evolution of Voice Systems and the Role of LLMs
The conversation with Bibbo Shu, PM Lead for Gemini Multimodeling and Project Astra, delves into the transformative journey of voice-based AI systems. She highlights that the current learning is focused on building "super fluid" communication systems, with LLMs being the key enabler for achieving "human-level dialogue systems." This core technology is expected to unlock a vast array of new products and services.
Bibbo Shu's Journey into Voice Technology
Shu's interest in voice technology began during her MBA at MIT Sloan, where she explored Human-Computer Interaction (HCI) at the MIT Media Lab. The concept of communicating with computers using the "entire body" – including voice and expressiveness – profoundly impacted her. This led her to join the original Google Assistant team, initially focused on voice search and actions. Her role as the original PM for Google Home was pivotal, marking her involvement in early "assisty" technologies. Her encounter with linguists and an audio class further solidified her passion for language, spoken dialogue, and its potential for creating understanding. She emphasizes that language, particularly spoken language, is a unique human mechanism.
Historical Milestones in Voice Recognition
The evolution of voice systems is traced back to early work at Bell Labs on number recognition. A significant breakthrough occurred around 2012 with Google research. Shu was involved in the period leading up to the Google Home product, working with teams pioneering "far-field speech recognition." This technology, combined with echo cancellation, allowed devices to hear users from a distance, a crucial step for products like Google Home. She recalls struggling to get users to speak commands into a phone's text box for music playback before Google Assistant existed, illustrating the user mental model challenge. The prototype for Google Home involved a tablet and a far-field microphone array, which proved highly effective in a home setting. Speech recognition accuracy also improved dramatically, with WER dropping from around 20% to 5-8% by 2015-2016, making systems usable.
Modalities, Spaces, and User Comfort
The discussion highlights how the modality and context of interaction influence user comfort. While users were hesitant to speak commands into a phone search box, they readily adopted voice commands with the Google Home prototype in their kitchen. This suggests a UX element tied to the physical space and the perceived purpose of the device. Voice, while powerful, has limitations, particularly in public spaces due to social norms. However, it excels in situations where hands-free operation is necessary, such as in the home or car, where pulling out a phone is inconvenient.
The LLM Revolution: From Brittle to Fluid Dialogue
A fundamental shift is occurring with LLM-based audio systems, moving beyond the "brittle" command-and-control era of home products. Previously, systems could handle simple back-and-forth for form-filling (e.g., "Set a timer for how long? Five minutes."), but any deviation would cause them to break. LLMs enable "free-flowing dialogue" where the AI "understands everything that you're saying." This is a "step change" that transforms AI from a command-and-control tool to a conversational partner capable of providing advice and engaging in uniquely human experiences.
Unexpected User Behaviors and the Evolution of Expectations
The evolution of AI has led to unexpected user behaviors. A user research example involved a user wanting Google Maps to be more interactive, allowing for dialogue about navigation choices (e.g., questioning an exit). This indicates a shift in user expectations: if an AI can understand anything, why can't it be a collaborative participant in tasks? This is a capability previously impossible but now achievable with advanced AI.
Deep Dive into LLM Advancements in Voice
Shu breaks down the impact of LLMs on voice technology into several key areas:
- Fluid Human-Level Dialogue: The goal is to achieve near-perfect human-level dialogue, addressing aspects like smooth interjections, the use of "ums" and "ahs" (backchanneling), and more natural expression.
- Expressive Speech and Tone: Modern audio models generate speech with tone, pitch, and inflection, making them sound expressive rather than robotic. This enhances communication and understanding.
- Tone and Emotion: The way something is said, including emphasis and emotion, is as important as the words themselves. AI is beginning to capture and convey these nuances.
- Translation: LLMs have revolutionized translation. Instead of a cascade of transcription, text translation, and speech synthesis, multimodal systems can perform direct audio-to-audio translation in real-time, preserving the speaker's voice. The live speech-to-speech translation in Google Meet is a prime example, where a Spanish speaker's voice is heard in English.
- Code Mixing and Language Adaptability: AI models can now handle code mixing, where users blend languages. The AI can understand and respond in the user's chosen language mix, adapting dynamically within a conversation. This is possible because native audio models process audio directly without needing explicit language identification steps.
Project Astra: A Situated Universal Assistant
Project Astra is described as a "situated assistant" that "sees the world with you." It's a research prototype aiming to be a "concept card for the universal assistant," integrating various capabilities. The core idea is dialogue with the camera on, allowing the AI to understand and discuss what it sees. A significant application demonstrated was assisting blind users with navigation. Astra was announced at Google I/O, showcasing Gemini's ability to understand the world and engage in dialogue about it.
Challenges in Building Multimodal AI
Integrating multiple modalities (audio, vision) and capabilities like memory presents significant challenges for developers. Shu recounts a "mind-blowing bug" where Astra reported "I can't see" due to a visual processing failure, but then continued to exhibit this behavior even after the system was fixed because it had "remembered" from a past conversation that it couldn't see. This highlights the difficulty in debugging and understanding AI memory. The integration of various critical capabilities for a universal assistant into a single model creates complex issues.
Advice for Developers
Shu offers advice for developers:
- Leverage Live API: The live API provides a strong foundation for coherent dialogue and search grounding, enabling the creation of voice agents for companies and new use cases in manufacturing and troubleshooting.
- Vision Capabilities: The vision capabilities of LLMs are now highly advanced, capable of identifying a wide range of objects and scenes with accuracy.
- Screen Sharing: Vision extends to the digital world through screen sharing, allowing AI to assist with on-screen tasks. An example is an AI helping to pronounce a difficult word encountered while reading an article.
Interruption and Latency: The Keys to Natural Dialogue
Interruption (barge-in) has always been important for dialogue systems, but historically, systems were slow and struggled to differentiate user interruptions from background noise. The critical factor for natural dialogue is latency. Innovations in low-latency LLM serving are crucial. The goal is to achieve under 500 milliseconds for human dialogue. Future research, like "mixtape" or "full duplex dialogue," aims for AI to continuously listen, even while speaking, mirroring human conversation. The "open mic" experience, where the mic is always on, is enabled by models that can distinguish intended input from background noise.
The Blurring Lines: Human vs. AI Interaction
The increasing human-likeness of AI raises ethical questions about distinguishing between human and AI interactions. While the goal is not to deceive, the ability of AI to communicate at a human level means it could be mistaken for a human. Shu emphasizes that the goal is to create a conversational AI that can communicate effectively and help users, not to fool them. The user expectation is shifting, and AI is beginning to understand and respond to user emotions, a capability that was absent in older systems.
Future of Work and Communication
- Automation for Learning: Shu uses Gemini as an automation system for learning, feeding it complex reports and then discussing them.
- Voice vs. Text with Gemini: She uses both voice and text with Gemini, often using deep research to create PDFs and then discussing them with Astra for a combination of deep knowledge and conversational interaction.
- UK Tax Questions: Her last query to Gemini was about complex UK tax changes, which she found helpful after struggling with a tax accountant.
- Coding and Tool Use: A significant new capability is enabling AI to "code and use tools" on behalf of the user, a development that has become more possible in the last year.
- Next Problems to Solve:
- Audio/Vision: Achieving human-level dialogue and "supernatural superfluid conversation" that is fully duplex.
- Execution: Enabling assistants to actively "do stuff" for users, not just talk through them. This involves computer control, API integration, and agentic automation systems (e.g., booking a tennis court on a difficult website).
- Biggest Learning (Communication): The key learning is how to build systems that communicate "super fluidly" using LLMs, leading to human-level dialogue systems that will enable many new products.
- Impact on Work: AI is fundamentally changing daily work. Writing, for instance, has been transformed, with users now able to "spew out" ideas and then use AI to refine them into coherent documents. This represents a new way of communicating thoughts and expressing oneself.
The conversation concludes by emphasizing the exciting, albeit sometimes scary, future of AI and communication, driven by continuous evolution and innovation.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development