Running Gemma on Mac and Windows PCs with Ollama

Google for DevelopersAbout 4 min readApr 5, 2025Watch original
THE SUMMARYAI-generated

Ollama and Gemma 3: Running AI Models Locally

Key Concepts:

  • Ollama: A portable inference engine for running AI models locally.
  • Gemma 3: A family of open-source AI models by Google DeepMind, now available on Ollama.
  • Local Inference: Running AI models directly on a user's machine (CPU, GPU, NPU, TPU) without relying on external APIs or cloud services.
  • Multi-modality: The ability of a model to process different types of input, such as text and images.
  • VRAM: Video RAM, the memory on a graphics card.
  • Google Cloud Run: A serverless execution environment for containerized applications.

1. Introduction to Ollama

  • Ollama is presented as the easiest way to run AI models locally.
  • It's an open-source project with binaries available for download or build from source on GitHub.
  • Users can interact with models via the terminal, Ollama's API, or open-source UIs.
  • Tens of millions of developers use Ollama, with 150 million models downloaded from ollama.com.
  • Ollama is a portable inference engine that handles model scheduling and memory utilization across CPUs, GPUs, NPUs, and TPUs.
  • It supports Mac, Windows, and Linux environments.

2. Gemma 3 on Ollama

  • Gemma 3 is now available on Ollama, enabling users to run the models locally.
  • All four parameter sizes of Gemma 3 are available on ollama.com.

3. Demo: Running Gemma 3 (4B Model)

  • The demo showcases running the Gemma 3 4B model on an M4 MacBook Pro (M4 Max).
  • The command ollama run gemma3 is used to initiate the model.
  • The model responds quickly to text prompts, demonstrating fast local inference.
  • No API keys or accounts are required.

4. Multi-Modal Capabilities Demo

  • The demo highlights Gemma 3's multi-modal capabilities using image input.
  • The 4B model successfully identifies a cow on a beach in an image.
  • A more complex example involves interpreting a doctor's note with poor handwriting.
  • Gemma 3 accurately identifies the medication (antibiotic), dosage (500mg), and instructions (one capsule, three times a day for seven days).

5. Gemma 3 (1B and 27B Models)

  • The 1B parameter model offers even faster inference and requires only 2GB of VRAM, making it suitable for almost any machine.
  • Developers report that the 1B model's responses are comparable to much larger models.
  • The 27B model requires more VRAM but is still usable on medium to high-end MacBooks, GeForce cards, and AMD graphics cards.

6. Ollama's API and Ecosystem

  • The demos use the command-line tool, but developers can leverage Ollama's built-in API or the 14,000+ GitHub projects built on top of Ollama.

7. Ollama on Google Cloud Run

  • Ollama has partnered with Google Cloud Run to enable serverless deployments.
  • Users can acquire a GPU on Google Cloud Run and scale from 0 to hundreds of GPUs.
  • GPU cold starts on Cloud Run are fast, with drivers installed in under five seconds.

8. Conclusion

  • Gemma 3 is an incredible set of models, and Ollama provides an easy way to run them locally.
  • Ollama is available on Windows, Mac, and Linux.

Notable Quotes:

  • Michael Chiang: "Ollama is the easiest way to run AI models locally."
  • Jeffrey Morgan: "We're so excited that Gemma 3 is available on Ollama starting today."

Technical Terms Explained:

  • Inference Engine: A software component that executes AI models to generate predictions or outputs based on input data.
  • Parameter Size: Refers to the number of parameters in a model, which generally correlates with its complexity and performance.
  • Serverless: A cloud computing execution model where the cloud provider dynamically manages the allocation of machine resources.

Logical Connections:

The presentation flows logically from introducing Ollama to demonstrating its capabilities with Gemma 3. It starts with the basics of Ollama, then moves to specific examples with different sizes of the Gemma 3 model, highlighting both text and image processing. Finally, it expands on Ollama's deployment options, including Google Cloud Run.

Synthesis/Conclusion:

Ollama simplifies the process of running AI models locally, as demonstrated by its integration with Google's Gemma 3. The presentation highlights the ease of use, multi-modal capabilities, and scalability of Ollama, making it a valuable tool for developers and researchers. The availability of Gemma 3 on Ollama empowers users to experiment with and deploy AI models on their own hardware or in serverless environments.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.