A truly multilingual Gemma 3

Google for DevelopersAbout 3 min readApr 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Gemma 3: Google's multilingual language model.
  • Multilingual Performance: The ability of a model to perform well in multiple languages.
  • Instruction-Tuning: Fine-tuning a language model to follow instructions effectively.
  • Tokenizer: A component that breaks down text into smaller units (tokens) for processing.
  • Pre-training: Training a language model on a large dataset before fine-tuning.
  • Post-training: Further training a pre-trained model on a specific task or dataset.
  • SEA-LION: A Gemma-based model tailored for Southeast Asian languages.
  • Project Aquarium: A project to make curated datasets for Southeast Asian languages available to the public.

1. Introduction and Motivation:

  • Adi Mayrav Gilady, a product manager at Google, discusses the multilingual capabilities of Gemma 3.
  • 80% of Gemma users are from outside the US, and over 64% are from countries where English is not the primary language.
  • Developers want to reach global audiences in their native languages.
  • A Kaggle competition highlighted the community's need for multilingual and multicultural models, receiving over 7,000 entries.

2. Gemma 3's Multilingual Enhancements:

  • Gemma 3 focuses on enhanced multilingual performance across various languages and tasks.
  • It features 140 languages represented in pre-training and over 35 languages working out-of-the-box in instruction-tuned settings.
  • Multilingual benchmarks showed improvements across machine translation, reasoning, and cross-lingual transfer compared to previous versions.
  • Internal evaluations show Gemma's instruction-tuned performance is comparable to GPT 4.0 in several languages.

3. How Multilingual Capabilities Were Improved:

  • Tokenizer: Gemma 3 uses a new tokenizer from the Gemini tokenizer, which favors non-English text.
  • Pre-training Data:
    • Carefully curated data with attention to quality within and across languages.
    • The multilingual share of content in pre-training was doubled.
  • Post-training:
    • A curated set of instruction-following prompts and datasets from human and synthetic methods.
    • Emphasis on covering a wide variety of use cases, languages, and cultural contexts.
    • Ensuring each sample feels native and fluent.
  • Evaluation: Continuous evaluation on multilingual benchmarks and careful assessment of output quality.
  • Maintaining English language quality and other capabilities like reasoning and coding.

4. Example Use Case:

  • An example is given where Gemma identifies languages in a sign from Berlin and provides context about the sign's historical significance (Cold War era).
  • This highlights the importance of combining multilingual capabilities with other features like multimodal understanding, long context, and reasoning.

5. Potential Applications:

  • Businesses (e.g., financial institutions, healthcare providers) can build chatbot support platforms in hundreds of languages.
  • Startups can use Gemma to build product integration tools supporting multiple languages.
  • Researchers can consume and analyze information in one language and produce insights in another.
  • Educators can bridge gaps with students in various locations.
  • Gemma can be used to preserve cultural contexts, especially for lesser-known languages.

6. SEA-LION Model: A Case Study:

  • SEA-LION is a collaboration between Google and AI Singapore, tailored for Southeast Asian languages.
  • Southeast Asia has over 1,000 native languages.
  • The model was continually pre-trained and instruction-tuned, focusing on these languages.
  • SEA-LION performs better than similarly sized models on Southeast Asian question benchmarks.
  • High-quality training data was curated by linguists and native speakers.
  • Project Aquarium will make this curated data available for public use and contribution.

7. Tips for Developers Fine-Tuning Gemma 3:

  • Evaluation: Having a quality evaluation set is crucial. Translated datasets are better than none.
  • Post-training Gains: Gemma's tokenizer and pre-training can lead to high gains in post-training. A post-trained model for a specific language can outperform a monolingually trained model.
  • Instruction-Tuning Samples: Even a small number of instruction-tuning samples (e.g., 40) can significantly improve instruction-following ability, even in unseen languages.

8. Conclusion:

  • Gemma 3 is live and available for developers to use via Google AI for Developers and partner platforms.
  • The model's enhanced multilingual capabilities, combined with other features, make it a powerful tool for diverse global applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.