How to Use Google Gemma 3.0 Multimodal Models FREE - 3 WAYS with Ollama, HuggingFace, Google Studio

DS-AI with Khanh VyAbout 4 min readMar 17, 2025Watch original
THE SUMMARYAI-generated

Jemma 3 Multimodal Models: A Tutorial

Key Concepts:

  • Jemma 3 (3B, 4B, 12B, 27B): Google's new family of open-source multimodal models.
  • AI Studio: Google's platform for experimenting with AI models.
  • Hugging Face: A platform for sharing and accessing pre-trained models.
  • Google Colab: A cloud-based platform for running Python code, especially for machine learning.
  • A100 GPU: A high-performance GPU commonly used for deep learning tasks.
  • Hugging Face Token: A credential used to access gated models on Hugging Face.
  • Pipeline (Image-to-Text): A machine learning pipeline that takes an image as input and generates text as output.
  • Ollama: A tool for running large language models locally.
  • Gated Model: A model that requires permission to access.

1. Experimenting with Jemma 3 via AI Studio

  • Access: Jemma 3 can be accessed directly through AI Studio (ai.google.com).
  • Model Availability: Initially, only Jemma 3 27B was accessible.
  • Text-Based Interaction: The tutorial demonstrates basic text prompting, asking for restaurant recommendations in Chicago. The model provides suggestions with cuisine types (Italian, Mexican, Steak House) and specific restaurant names.
  • Image Input Issue: The initial attempt to use the image input feature in AI Studio failed, with a message indicating that the current model doesn't support images. The presenter suggests this functionality may be updated later.

2. Using Hugging Face and Google Colab

  • Model Choice: Jemma 3 4B is used due to its lower computational requirements compared to the 27B version. The 27B version caused issues on a Google Colab A100 GPU.
  • Accessing Gated Model: Access to the Jemma 3 4B model on Hugging Face requires requesting and being granted access. This involves agreeing to terms and providing personal information (name, organization).
  • Hugging Face Token: A Hugging Face token is required to authenticate access to the model from Google Colab. The token needs to be created with "read" access for public gated repos. The presenter emphasizes the importance of securely storing the token after creation.
  • Google Colab Setup:
    • Enable a GPU runtime (A100 recommended) for faster processing.
    • Install necessary libraries: pip install git+https://github.com/google/generative-ai and pip install huggingface_hub.
    • Log in to Hugging Face using the generated token.
  • Image-to-Text Pipeline:
    • A pipeline is created for image-to-text tasks, enabling interaction with images.
    • The presenter sets up a message with a system role defining the model as a "helpful assistant." This role can be customized based on the desired behavior.
    • An image URL of multiple dogs is used as input.
    • The prompt asks the model to count the number of dogs and identify their breeds.
  • Results and Analysis:
    • The model attempts to count the dogs, estimating "20 to 25," which is inaccurate (actual count is 19). The presenter notes that other models (Cro, GPT) also struggled with this task.
    • The model identifies some breeds correctly (e.g., Border Collie, Golden Retriever) but also makes some inaccurate suggestions.
    • A second example uses an image of candy with a turtle design. The model correctly identifies the animal as a turtle, reasoning based on the shell shape, head, and legs.

3. Running Jemma 3 Locally with Ollama

  • Ollama Introduction: Ollama allows running large language models locally.
  • Model Selection: The presenter recommends using Jemma 3 4B for local execution, especially on systems with limited resources.
  • Pulling the Model: The command ollama pull jemma-3-4b is used to download the model.
  • Running the Model: The command ollama run jemma-3-4b starts the model.
  • Text-Based Interaction: The presenter asks for restaurant recommendations in Chicago with steakhouses.
  • Results: The model provides restaurant suggestions, including URLs and price ranges. The presenter praises the model's ability to provide accurate and relevant information, contrasting it with earlier experiences with GPT-3.

4. Conclusion

  • Jemma 3 offers various use cases, including chatbots and AI agents.
  • The ability to chat with images and videos (as claimed by Google) opens up new possibilities.
  • The presenter encourages viewers to provide feedback and questions.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.