Gemma 4 and the AI Edge Gallery: On-Device AI Gets an Upgrade

By Google for Developers

Share:

Key Concepts

  • Gemma 4: The latest iteration of Google’s open-model family, featuring sizes ranging from 2B to 31B parameters.
  • AI Edge Gallery: An on-device AI showcase application that allows users to run models locally without internet connectivity.
  • Agent Skills: A framework within the app that extends model capabilities by providing external tools and specific instructions.
  • MCP (Model Context Protocol): A standard for connecting AI models to external data sources, tools, and ecosystems.
  • On-Device Inference: Running AI models directly on hardware (phones/laptops) rather than in the cloud, ensuring privacy and offline functionality.
  • Structured Output: The ability of the model to format data into specific structures like JSON or bulleted lists.

1. Gemma 4 Overview

Gemma 4 is the latest generation of Google’s AI models, designed for diverse hardware environments:

  • Small Models (2B, 4B): Optimized for mobile devices and on-device execution.
  • Large Models (26B, 31B): Targeted at laptops, desktops, and servers.
  • Performance: The 2B model is noted for matching the performance of previous 27B dense models, demonstrating significant efficiency gains.

2. The AI Edge Gallery App

The AI Edge Gallery serves as a sandbox for developers and users to interact with Gemma 4.

  • Accessibility: Available on both the App Store and Play Store.
  • Performance: Within one month of launch, the app surpassed 5 million downloads.
  • Key Benefits: Eliminates token limits, reduces costs, and functions entirely offline, making it ideal for travel or low-connectivity environments.

3. Agent Skills and Tool-Calling

The "Agent Skills" feature allows the model to perform tasks beyond its base training window:

  • Web Fetch/Wikipedia: By hooking the model to external data sources, it can answer questions about current events or information outside its training data.
  • Rich Content Rendering: The model can output HTML, JavaScript, and interactive UI elements, enabling mini-games or interactive maps directly within the chat interface.
  • Implementation: Users can load skills via URL or local files. The app includes a community-driven repository where users can share and discover new skills.

4. Future Integrations: Model Context Protocol (MCP)

Google is introducing MCP integration to bridge the gap between local models and broader digital ecosystems:

  • Experimental Rollout: Starting on Android and expanding to iOS.
  • Use Cases:
    • Web Summarization: Linking to web-fetch tools to summarize blog posts or articles.
    • Game Integration: Using natural language to control game characters or manage inventories.
    • Server-Side Data: Accessing local server datasets to generate reports and summaries while keeping data private.

5. Multimodal Capabilities and Practical Applications

Gemma 4 supports multimodal inputs (audio and image), enabling real-world utility:

  • Translation & Context: Users can photograph menus in foreign languages; the model translates the text and explains cultural nuances.
  • Data Extraction: Users can photograph bookshelves or documents and request the model to output the information in a structured JSON format.
  • Voice Memos: The model can process spoken lists (e.g., shopping lists) and instantly format them into structured bullet points.

6. Development and Community

  • Open Source: The AI Edge Gallery is fully open-sourced on GitHub, allowing developers to inspect the code and build custom skills.
  • Persistent History: A new feature allows for long-term chat sessions, enabling users to pick up complex planning tasks (like travel itineraries) across different days.
  • Speed: Utilizing "fast preview" technology, the app achieves speeds of over 3,000 tokens per second on modern GPUs.

Notable Quotes

  • Olivier Lacombe: "I can see a future where maybe next year, our smaller size are competing with our 31B and it creates a whole new narrative of what you can actually do at the Edge."
  • Alice Zheng: "We want to inspire people. We want to see what people are using Gemma to solve the real problems and provide real value in their life."

Synthesis

The transition toward on-device AI, led by Gemma 4 and the AI Edge Gallery, marks a shift from cloud-dependent AI to localized, private, and highly capable personal assistants. By leveraging Agent Skills and the Model Context Protocol, Google is transforming static language models into active agents capable of interacting with apps, games, and external data, all while maintaining the privacy and speed benefits of local execution. The open-source nature of the project encourages community-led innovation, as seen in specialized use cases like sub-Saharan language support.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video