Build an AI Voice Translator: Keep Your Voice in Any Language! (Python + Gradio Tutorial)
By AssemblyAI
Key Concepts
- Voice Translator App: An application that records speech, translates it into multiple languages, and generates the translated speech using the user's own cloned voice.
- Gradio: A Python library for quickly building customizable machine learning web interfaces, used here for the app's UI.
- AssemblyAI: An AI service used for high-accuracy speech-to-text transcription, converting spoken English into text.
- Python
translatemodule: A Python library for text translation, utilizing various providers (e.g.,MyMemoryas a free default). - Eleven Labs: An AI voice technology platform used for text-to-speech generation and advanced voice cloning.
- Voice Cloning: The process of creating a synthetic voice that sounds like a specific person.
- Instant Voice Cloning: Requires only about one minute of audio.
- Professional Voice Cloning: A paid feature requiring at least 30 minutes of audio (optimal 3 hours) for a high-fidelity clone.
- Multilingual Model (
eleven_multilingual_v2): An Eleven Labs model capable of generating speech in multiple languages. - API Keys: Credentials required to access and authenticate with AssemblyAI and Eleven Labs services.
gradio.Interfacevs.gradio.Blocks: Two methods for building Gradio apps;Interfaceis simpler for direct input/output,Blocksoffers more layout customization.Pathlib: A Python module for object-oriented filesystem paths, used here to ensure Gradio correctly plays generated audio files.
Introduction to the Voice Translator App
The video introduces a voice translator application that records the user's speech in English, translates it into various languages, and then generates the translated speech using the user's own voice. The creator describes this capability as "mind-blowing" and "eerie" due to the experience of hearing oneself speak languages one doesn't know.
As an initial demonstration, the speaker recorded the phrase: "Hello, good morning. My name is Mr. and I'm excited to share this app that I made with everyone else." This phrase was then translated and generated in six different languages, including Turkish, Russian, and Japanese, showcasing the app's core functionality.
Building the Voice Translator App: Technologies and Architecture
The application is built using three primary technologies:
- Gradio: For creating the user interface.
- AssemblyAI: For transcribing the initial English speech into text.
- Python
translatemodule: For translating the transcribed text into target languages. - Eleven Labs: For generating audio of the translated text using the user's cloned voice.
The overall process involves a main voice_to_voice function that orchestrates calls to audio_transcription (AssemblyAI), text_translation (Python translate), and text_to_speech (Eleven Labs).
Setting Up the Gradio Layout
The tutorial focuses on building a simplified Gradio interface using gradio.Interface for ease of understanding, though a more complex version with additional features (like waveform display, download buttons, and text output) is available on the GitHub repository.
Gradio Interface Components:
- Input: A
gradio.Audiocomponent configured to use themicrophoneas the source, returning thefile_pathof the recorded audio. - Outputs: Three
gradio.Audiocomponents, labeled "Spanish," "Turkish," and "Japanese," to display the generated translated audio.
Running Gradio Apps: The speaker notes two ways to run Gradio apps:
python simple.py: Runs the app once.gradio app_name.py: Enables hot-reloading, updating the app in the browser upon saving changes to the code.
Implementing Audio Transcription with AssemblyAI
The audio_transcription function handles converting the recorded audio into text:
- Import:
import assemblyai as aai. - API Key: Set the AssemblyAI API key.
- Transcriber: Initialize
aai.Transcriber(). - Transcription: Call
transcribe()on theaudio_filepath. The Python SDK for AssemblyAI will hold execution until the transcription is complete. - Error Handling: The
transcription_responseobject is checked forassemblyai.TranscriptionStatus.error. If an error occurs, agradio.Erroris raised with the error message from AssemblyAI. - Text Extraction: If successful, the transcribed English text is extracted from
transcription_response.text.
Translating Text with Python's translate Module
The text_translation function takes the English text and translates it into the desired target languages:
- Import:
from translate import Translator. - Translator Instances: Create separate
Translatorinstances for each target language, specifyingfrom_lang="en"and the respectiveto_lang(e.g.,"es"for Spanish,"tr"for Turkish,"ja"for Japanese). Thetranslatemodule usesMyMemoryas its default free provider, which is deemed "good enough" for personal use, though paid options like Microsoft Translate are available for potentially better context-aware translations. - Translation: Call the
translate()method on eachTranslatorinstance, passing the English text. - Return: The function returns the translated texts for Spanish, Turkish, and Japanese. The speaker acknowledges this "hacking a solution together" for multiple languages could be optimized for efficiency.
Generating Audio with Eleven Labs
The text_to_speech function converts the translated text into audio using the user's cloned voice:
Prerequisites for Eleven Labs:
- Installation: Install
elevenlabsandpython-dotenv. - API Key: Obtain the Eleven Labs API key from
My Account->Profile->API Keyon the Eleven Labs dashboard. - Voice Cloning: To generate audio in the user's own voice, the voice must first be cloned on the Eleven Labs dashboard (
Voicessection).- Instant Voice Cloning: Requires only one minute of audio.
- Professional Voice Cloning: A paid feature requiring a subscription. It needs at least 30 minutes of audio for a good clone (the speaker provided ~33 minutes for the demo) and is optimal with around 3 hours of audio for a "flawless clone."
Eleven Labs Integration Steps:
- Import: Import
ElevenLabsClientandVoiceSettings. - API Key: Manually specify the Eleven Labs API key.
- Text-to-Speech Conversion: Call
client.text_to_speech.convert().voice_id: The ID of the cloned voice (copied from the Eleven Labs dashboard) is specified.model_id: Set to"eleven_multilingual_v2"to support multiple languages.VoiceSettings: Configured withstability=0.5,similarity_boost=0.8, andstyle_exaggeration=0.5based on the speaker's experimentation for optimal results.
- Save Audio: The generated audio stream is saved to a unique
.mp3file usinguuidfor filename generation. - Return Path: The function returns the
file_pathof the saved audio. - Gradio Compatibility: An extra step is required to convert the string
file_pathinto aPathlibobject (from pathlib import Path) before returning it to Gradio, as Gradio requires this format for proper audio playback.
Final Integration and Demo
The voice_to_voice main function integrates all these steps:
- It receives the
audio_filepath from the Gradio input. - Calls
audio_transcriptionto get the English text. - Calls
text_translationto get the Spanish, Turkish, and Japanese texts. - Calls
text_to_speechthree times, once for each translated text, to generate the respective audio files in the cloned voice. - Converts the returned audio file paths to
Pathlibobjects. - Finally, it returns the three
Pathlibobjects (for Spanish, Turkish, and Japanese audio) to the Gradio interface, which automatically displays them in the designated output audio components.
During the final demo, the speaker recorded: "Hello, it is a beautiful day today, but I'm a little bit cold because the AC in this room is blowing really hard." The app successfully processed this, generating the translated audio in the speaker's voice, which was described as "super fun."
Conclusion and Future Ideas
The tutorial concludes by reiterating the ease of building such a powerful application with Gradio and the integrated AI services. The speaker mentions that a more complex Gradio interface, similar to the initial demo with visual waveforms and download options, is available on GitHub.
Potential Applications and Ideas:
- Samsung's Voice-to-Voice Translation: The speaker notes Samsung's recent implementation of real-time voice translation during phone calls, where voices are translated (though not using the original speaker's voice).
- Personalized Communication: Sending WhatsApp voicemails to friends in their native language using one's own voice, which would be a "quite a bit" surprising and engaging experience.
- Language Practice: Using the app to practice speaking a new language by imitating one's own voice in the translated audio, potentially making it easier to mimic pronunciation.
The speaker encourages viewers to share their own ideas for using this technology in the comments, suggesting that a future tutorial could even be based on a community idea. A related resource, Smita's video on building an AI voice bot (a ChatGPT bot with voice interaction), is also recommended.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development