Integrating the Gemma model family in Transformers

Google for DevelopersAbout 4 min readApr 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Transformers library as a source of truth
  • Model integration best practices
  • Device-agnostic code
  • Model maintenance
  • Common API and features
  • Standardization, simplification, and reproducibility
  • Model comparison (Llama standard)
  • Hugging Face repository
  • Attention mask complexities
  • Reproducible snippet
  • Integration tests
  • Integration parties (you, us, our side)
  • CLI tool
  • Precision issues (bfloat16)
  • Rotary Position Embedding (RoPE)
  • Head resizing

Why Integrate Models into Transformers?

The speaker, Arthur Zucker from Hugging Face, explains the benefits of integrating models into the Transformers library. The primary reason is that Transformers aims to be the "source of truth," ensuring code compatibility across various AI frameworks. This compatibility extends to:

  • Inference: Transformers serves as the backend for state-of-the-art inference libraries like TGI and vLLM.
  • Training: Libraries like Unsloth, Axolotl, Liger, DeepSpeed, and Accelerate are fully compatible with Transformers.
  • Quantization: Transformers explores various open-source quantization methods.
  • Research: Collaboration with teams like NVIDIA on techniques like kvpress and special generating methods.

A key advantage is the device-agnostic nature of the code, allowing models to run on CPUs, GPUs, TPUs, and other devices. Furthermore, Transformers provides long-term maintenance, ensuring that even older models like the original BERT remain functional and testable, enabling comparisons with newer models like modern BERT. The library also offers a common API (e.g., from_pretrained, tokenizer) and features, making it familiar to users. Integration into Transformers provides significant visibility, with around 50 million downloads per month, exposing models to a large audience actively seeking new techniques. The team prioritizes standardization, simplification, and reproducibility, which, while challenging, ensures consistent performance.

Requirements for Model Integration

To integrate a model into Transformers, several key components are needed:

  1. Model Comparison: Comparing the model to a standard like Llama helps identify differences. Inheritance can be used to highlight specific layer variations. For example, Gemma 3 inherits from Gemma 2, making it easy to pinpoint the unique layers (e.g., KQ norm, K norm). A concise summary of the model's unique features is also beneficial.
  2. Weights: Providing open-source and open-weight models is crucial. Hugging Face repositories are recommended for easy collaboration and access control.
  3. Helpers: Explanations for complex aspects, such as the attention mask, are essential. ASCII diagrams can illustrate how image tokens should attend to each other or how the mask should behave in sliding window scenarios. Examples like the PaliGemma attention mask, which differs significantly from standard masks, highlight the importance of clear explanations.
  4. Input/Output Examples: Providing input and output examples, including token IDs and tokens, helps understand the model's specific behavior. This is especially important because different tokenization frameworks (e.g., Tiktoken, SentencePiece, Tokenizers) may be used.
  5. Reproducible Snippet: A reproducible code snippet is standard and crucial for creating integration tests that are maintained over time. These tests ensure that the model functions correctly, even in edge cases like sliding windows.

Integration Process

There are three main ways a model can be integrated:

  1. External Contribution: The model creator (you or your organization) integrates the model. This allows control over the timeline but can be challenging due to unfamiliarity with the Transformers codebase and potential delays in code review.
  2. Collaboration: A collaborative effort between the model creator and the Transformers team. This is considered the best approach, allowing for direct communication and addressing specific questions. The Gemma integrations followed this model.
  3. Transformers Team Integration: The Transformers team integrates the model independently. This doesn't guarantee a day-zero integration due to resource constraints. A private fork of Transformers or an open PR with trust_remote_code=True may be used initially. The team strives to integrate models fully to ensure all features, such as batched inference, are available.

The Transformers CLI tool is essential for creating the necessary files for core features like from_pretrained and AutoConfig. It helps identify the closest existing model, simplifying the integration process.

Common Pitfalls and Technical Details

The speaker highlights several common issues encountered during model integration:

  • Precision Issues (bfloat16): Using bfloat16 with large integers can lead to significant errors. For example, when calculating rotary position embeddings (RoPE), multiplying float32 values with position IDs (which can be large) requires careful handling. Casting position IDs to bfloat16 prematurely can result in incorrect outputs.
  • Rotary Position Embedding (RoPE): License restrictions may force the use of alternative RoPE formulations (e.g., EleutherAI's version), which require rotating the query and key weights. This can lead to discrepancies and necessitate reshaping, transposing, and repeating weights, especially when query norm or K norm layers are involved.
  • Head Resizing: Resizing the language model head or embedding layer can disrupt the distribution probabilities of tokens, significantly affecting the model's output.

Conclusion

Integrating models into the Transformers library offers numerous benefits, including code compatibility, long-term maintenance, and increased visibility. However, the process requires careful attention to detail, including thorough model comparison, clear documentation, and adherence to best practices. Common pitfalls, such as precision issues with bfloat16 and complexities with RoPE implementations, must be addressed to ensure accurate and reliable model performance. The Transformers team provides tools and support to facilitate the integration process, but collaboration between model creators and the team is often the most effective approach.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.