Key Concepts
- Multilingual Conversational AI Agents
- Speech-to-Text (STT) / Automatic Speech Recognition (ASR)
- Large Language Models (LLMs)
- Text-to-Speech (TTS)
- Language Detection
- Function Calling / Tool Calling
- Voice Cloning and Royalties
- Low Latency
- Agent-to-Agent Transfers
- Safety Tooling (Watermarking, Moderation)
- Pronunciation Dictionaries
1. Introduction and 11 Labs Overview
- Thor and Paul from 11 Labs introduce themselves, focusing on developer experience related to conversational AI agents.
- 11 Labs specializes in speech technology, partnering with LLM providers for the "brain" of AI agents.
- Attendees are encouraged to scan a QR code for slides, resources, and a form to receive credits for using 11 Labs.
- The 11 Labs Devs Twitter account is recommended for API updates and developer-related news.
2. Multilingual Capabilities and Language Support
- The workshop focuses on building multilingual conversational AI agents.
- Attendees mention specific languages of interest: Portuguese (Brazilian accent), Spanish, Hungarian, Mandarin, and Hindi.
- 11 Labs is working on version 3 of its multilingual models, aiming to support up to 99 languages.
- Currently, Hindi and Tamil are supported, with plans to add more Indian languages.
3. Unexpected Use Case: Text-to-Bark
- 11 Labs launched a text-to-bark model, allowing users to generate sounds for dogs.
- The launch coincided with April Fools' Day, leading to initial skepticism.
- The model is part of 11 Labs' sound effects model, which can generate various audio samples (e.g., truck reversing, drum cowbell).
- A "sound board" drum machine was created using these sound effects.
4. Conversational AI Agent Pipeline
- The core pipeline involves:
- User speech input.
- Speech-to-text transcription.
- LLM processing (acting as the agent's brain).
- Text-to-speech output.
- System tools are integrated, including language detection and function calling.
- 11 Labs prioritizes a text-based pipeline for better monitoring and understanding of conversations, contrasting with sound token to sound token approaches.
- Models are deployed close to each other to minimize latency.
5. Speech-to-Text (ASR) Model
- 11 Labs' ASR model supports 99 languages with benchmark-leading performance.
- Features include word-level timestamps, speaker diarization, and audio event tagging (e.g., "coughing," "laughing").
- Structured API responses provide easy access to these features.
- Example: A conference call transcript demonstrates speaker identification and word-level timestamps.
6. Telegram Transcription Bot Example
- A Telegram bot was created to transcribe voice messages and videos.
- The bot automatically identifies the language and provides a transcript.
- Examples:
- Transcribing a message in "Singlish" (Singaporean English).
- Transcribing a Scottish accent.
- Transcribing speech with poor audio quality and background noise.
- The bot showcases the model's ability to handle various accents and audio conditions.
7. LLM Integration and Customization
- 11 Labs partners with leading LLM providers but allows users to integrate custom fine-tuned models.
- Custom LLMs require an OpenAI API-compatible endpoint (e.g., deployed on Google Vertex AI).
8. Text-to-Speech (TTS) and Voice Library
- 11 Labs offers a library of over 5,000 voices.
- Users can filter voices by language, accent, gender, and age.
- Voice cloning is supported, allowing users to create digital versions of their own voices.
- A marketplace exists where voice actors can publish their voices and earn royalties.
- Example: The presenter's voice is available in the voice library, generating royalties for him when used.
9. Conversational AI Agent Configuration
- Agents can be configured via a dashboard or the API.
- Dashboard configuration includes:
- Choosing an LLM provider or using a custom LLM.
- Adding a knowledge base (documents, website references, RAG).
- Configuring tools (function calling).
- Enabling system tools (language detection).
- Example: A Singapore-based agent supports English, Mandarin Chinese, Malay, and Tamil.
10. Language Detection System Tool
- The language detection tool can automatically identify and switch between languages.
- Two modes:
- Automatic detection based on the user's speech.
- Explicit language switching based on user requests (e.g., "Can we switch to Hindi?").
- The tool assigns a confidence score to each language, determining the likelihood of it being spoken.
11. Agent Configuration and Customization
- Agents can be configured in the dashboard or via the API.
- API configuration is suitable for building marketplaces where agents are configured on behalf of others.
- An MCP server is available for cloud desktop environments, allowing users to set up agents using natural language commands.
12. Q&A Highlights
- Language Switching: The ASR model identifies the spoken language and switches to the corresponding voice configured in the agent settings.
- Function Calling: Agents can call webhooks to integrate with external systems (e.g., CRM, scheduling tools).
- Low Latency: Using smaller LLMs (e.g., Gemini Flash) and 11 Labs' flash voice models can reduce latency.
- Cost per Minute: Pricing is based on call minutes and varies depending on the pricing tier.
- Multi-Agent Configuration: Agent-to-agent transfers allow for routing conversations to different agents based on the task.
- Handling Latency: The agent can provide conversational updates while waiting for tool responses, with configurable timeout settings.
- Safety: 11 Labs employs watermarking, moderation, and voice capture to prevent misuse of the technology.
- Mixed Languages: The ASR model's performance may degrade with more than two intermixed languages.
- Lip Sync: 11 Labs primarily partners with companies like Hedra and HeyGen for avatar-related technologies.
- Custom Vocabulary: Pronunciation dictionaries can be used to ensure correct pronunciation of acronyms and specific words in TTS. System prompts can be used to normalize acronyms in STT.
13. Synthesis/Conclusion
The workshop provides a detailed overview of building multilingual conversational AI agents using 11 Labs' platform. It covers the core components of the pipeline (STT, LLM, TTS), highlights the platform's multilingual capabilities, and demonstrates practical examples of agent configuration and customization. The Q&A session addresses key concerns such as latency, cost, safety, and handling complex linguistic scenarios. The main takeaway is that 11 Labs offers a comprehensive suite of tools and resources for developers to create sophisticated and engaging conversational AI experiences in multiple languages.
AI summaries can miss context or contain errors. Check important details against the original video.





