Build an AI Voice Translator: Keep Your Voice in Any Language! (Python + Gradio Tutorial)

By AssemblyAI

Share:

Key Concepts

  • Voice Translator App: An application that records speech, translates it into multiple languages, and generates the translated speech using the user's own cloned voice.
  • Gradio: A Python library for quickly building customizable machine learning web interfaces, used here for the app's UI.
  • AssemblyAI: An AI service used for high-accuracy speech-to-text transcription, converting spoken English into text.
  • Python translate module: A Python library for text translation, utilizing various providers (e.g., MyMemory as a free default).
  • Eleven Labs: An AI voice technology platform used for text-to-speech generation and advanced voice cloning.
  • Voice Cloning: The process of creating a synthetic voice that sounds like a specific person.
    • Instant Voice Cloning: Requires only about one minute of audio.
    • Professional Voice Cloning: A paid feature requiring at least 30 minutes of audio (optimal 3 hours) for a high-fidelity clone.
  • Multilingual Model (eleven_multilingual_v2): An Eleven Labs model capable of generating speech in multiple languages.
  • API Keys: Credentials required to access and authenticate with AssemblyAI and Eleven Labs services.
  • gradio.Interface vs. gradio.Blocks: Two methods for building Gradio apps; Interface is simpler for direct input/output, Blocks offers more layout customization.
  • Pathlib: A Python module for object-oriented filesystem paths, used here to ensure Gradio correctly plays generated audio files.

Introduction to the Voice Translator App

The video introduces a voice translator application that records the user's speech in English, translates it into various languages, and then generates the translated speech using the user's own voice. The creator describes this capability as "mind-blowing" and "eerie" due to the experience of hearing oneself speak languages one doesn't know.

As an initial demonstration, the speaker recorded the phrase: "Hello, good morning. My name is Mr. and I'm excited to share this app that I made with everyone else." This phrase was then translated and generated in six different languages, including Turkish, Russian, and Japanese, showcasing the app's core functionality.


Building the Voice Translator App: Technologies and Architecture

The application is built using three primary technologies:

  1. Gradio: For creating the user interface.
  2. AssemblyAI: For transcribing the initial English speech into text.
  3. Python translate module: For translating the transcribed text into target languages.
  4. Eleven Labs: For generating audio of the translated text using the user's cloned voice.

The overall process involves a main voice_to_voice function that orchestrates calls to audio_transcription (AssemblyAI), text_translation (Python translate), and text_to_speech (Eleven Labs).


Setting Up the Gradio Layout

The tutorial focuses on building a simplified Gradio interface using gradio.Interface for ease of understanding, though a more complex version with additional features (like waveform display, download buttons, and text output) is available on the GitHub repository.

Gradio Interface Components:

  • Input: A gradio.Audio component configured to use the microphone as the source, returning the file_path of the recorded audio.
  • Outputs: Three gradio.Audio components, labeled "Spanish," "Turkish," and "Japanese," to display the generated translated audio.

Running Gradio Apps: The speaker notes two ways to run Gradio apps:

  • python simple.py: Runs the app once.
  • gradio app_name.py: Enables hot-reloading, updating the app in the browser upon saving changes to the code.

Implementing Audio Transcription with AssemblyAI

The audio_transcription function handles converting the recorded audio into text:

  1. Import: import assemblyai as aai.
  2. API Key: Set the AssemblyAI API key.
  3. Transcriber: Initialize aai.Transcriber().
  4. Transcription: Call transcribe() on the audio_file path. The Python SDK for AssemblyAI will hold execution until the transcription is complete.
  5. Error Handling: The transcription_response object is checked for assemblyai.TranscriptionStatus.error. If an error occurs, a gradio.Error is raised with the error message from AssemblyAI.
  6. Text Extraction: If successful, the transcribed English text is extracted from transcription_response.text.

Translating Text with Python's translate Module

The text_translation function takes the English text and translates it into the desired target languages:

  1. Import: from translate import Translator.
  2. Translator Instances: Create separate Translator instances for each target language, specifying from_lang="en" and the respective to_lang (e.g., "es" for Spanish, "tr" for Turkish, "ja" for Japanese). The translate module uses MyMemory as its default free provider, which is deemed "good enough" for personal use, though paid options like Microsoft Translate are available for potentially better context-aware translations.
  3. Translation: Call the translate() method on each Translator instance, passing the English text.
  4. Return: The function returns the translated texts for Spanish, Turkish, and Japanese. The speaker acknowledges this "hacking a solution together" for multiple languages could be optimized for efficiency.

Generating Audio with Eleven Labs

The text_to_speech function converts the translated text into audio using the user's cloned voice:

Prerequisites for Eleven Labs:

  1. Installation: Install elevenlabs and python-dotenv.
  2. API Key: Obtain the Eleven Labs API key from My Account -> Profile -> API Key on the Eleven Labs dashboard.
  3. Voice Cloning: To generate audio in the user's own voice, the voice must first be cloned on the Eleven Labs dashboard (Voices section).
    • Instant Voice Cloning: Requires only one minute of audio.
    • Professional Voice Cloning: A paid feature requiring a subscription. It needs at least 30 minutes of audio for a good clone (the speaker provided ~33 minutes for the demo) and is optimal with around 3 hours of audio for a "flawless clone."

Eleven Labs Integration Steps:

  1. Import: Import ElevenLabsClient and VoiceSettings.
  2. API Key: Manually specify the Eleven Labs API key.
  3. Text-to-Speech Conversion: Call client.text_to_speech.convert().
    • voice_id: The ID of the cloned voice (copied from the Eleven Labs dashboard) is specified.
    • model_id: Set to "eleven_multilingual_v2" to support multiple languages.
    • VoiceSettings: Configured with stability=0.5, similarity_boost=0.8, and style_exaggeration=0.5 based on the speaker's experimentation for optimal results.
  4. Save Audio: The generated audio stream is saved to a unique .mp3 file using uuid for filename generation.
  5. Return Path: The function returns the file_path of the saved audio.
  6. Gradio Compatibility: An extra step is required to convert the string file_path into a Pathlib object (from pathlib import Path) before returning it to Gradio, as Gradio requires this format for proper audio playback.

Final Integration and Demo

The voice_to_voice main function integrates all these steps:

  1. It receives the audio_file path from the Gradio input.
  2. Calls audio_transcription to get the English text.
  3. Calls text_translation to get the Spanish, Turkish, and Japanese texts.
  4. Calls text_to_speech three times, once for each translated text, to generate the respective audio files in the cloned voice.
  5. Converts the returned audio file paths to Pathlib objects.
  6. Finally, it returns the three Pathlib objects (for Spanish, Turkish, and Japanese audio) to the Gradio interface, which automatically displays them in the designated output audio components.

During the final demo, the speaker recorded: "Hello, it is a beautiful day today, but I'm a little bit cold because the AC in this room is blowing really hard." The app successfully processed this, generating the translated audio in the speaker's voice, which was described as "super fun."


Conclusion and Future Ideas

The tutorial concludes by reiterating the ease of building such a powerful application with Gradio and the integrated AI services. The speaker mentions that a more complex Gradio interface, similar to the initial demo with visual waveforms and download options, is available on GitHub.

Potential Applications and Ideas:

  • Samsung's Voice-to-Voice Translation: The speaker notes Samsung's recent implementation of real-time voice translation during phone calls, where voices are translated (though not using the original speaker's voice).
  • Personalized Communication: Sending WhatsApp voicemails to friends in their native language using one's own voice, which would be a "quite a bit" surprising and engaging experience.
  • Language Practice: Using the app to practice speaking a new language by imitating one's own voice in the translated audio, potentially making it easier to mimic pronunciation.

The speaker encourages viewers to share their own ideas for using this technology in the comments, suggesting that a future tutorial could even be based on a community idea. A related resource, Smita's video on building an AI voice bot (a ChatGPT bot with voice interaction), is also recommended.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video