Gemma on mobile and web. Best and worst practices

Google for DevelopersAbout 5 min readApr 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • On-device AI: Running AI models directly on mobile and web devices instead of relying on cloud-based APIs.
  • System GenAI: LLMs baked directly into platforms like Android (Gemini Nano, AICore) and iOS (Apple Intelligence).
  • In-app AI: Bundling an AI model within an application, either as part of the APK or downloaded after installation.
  • Gemma 3 1B: A specific, smaller LLM model (1 billion parameters) optimized for on-device use.
  • Quantization: A technique to reduce the size of a model by using lower-precision numbers.
  • RAG (Retrieval Augmented Generation): A method to enhance LLM performance by retrieving relevant information from a knowledge base.
  • Embeddings: Vector representations of the intent of data, used in RAG for efficient information retrieval.
  • MediaPipe LLM Inference API: An API for running Gemma on Android, iOS, and web with cross-platform support.
  • Few-shot prompting: Providing examples of inputs and expected outputs in the prompt to guide the model.
  • Fine-tuning: Customizing a pre-trained model on a specific dataset to improve its performance on a particular task.

Why Use AI on Device?

Mark Sherwood outlines several key reasons for using AI on devices instead of relying solely on cloud-based solutions:

  • Cost: Eliminates cloud bills and enables different business models (free apps, freemium tiers, single payment apps) by removing variable costs associated with each user interaction.
  • Privacy: Keeps sensitive data on the device, ensuring it never leaves or is end-to-end encrypted, crucial for applications dealing with personal information.
  • Offline Availability: Allows apps to function without an internet connection, benefiting users in areas with poor connectivity or during travel.
  • Latency: Smaller on-device models can provide faster response times compared to cloud-based models, especially in areas with poor internet connections.

System GenAI vs. In-App AI

The presentation contrasts two approaches to using AI on devices:

  • System GenAI (e.g., Gemini Nano, Apple Intelligence):
    • Pros: Largest, most capable models; pre-downloaded for ease of use; best quality.
    • Cons: Limited flexibility and customization; access restricted to platform-provided UI elements, no open API access.
  • In-App AI (e.g., Gemma):
    • Pros: Full customization; complete control over UI and fine-tuning; faster performance due to smaller model size; wider device reach.
    • Cons: Requires bundling or downloading the model; potentially lower quality compared to larger system models.

Gemma 3 1B: An In-App Model

Gemma 3 1B is highlighted as an excellent in-app model due to its:

  • Size: Quantized size of only 529 MB (under 400 MB when zipped), making it easy to download over the air.
  • Device Reach: Can run on devices with 4 GB of RAM or more, covering virtually all devices in use.
  • Speed: Achieves 2,500 tokens per second prefill and 56 tokens per second decode speed, providing a near-instant experience.
  • Suitability for On-Device Use Cases: Well-suited for processing app context and generating short snippets.

Use Cases for Gemma in Applications

The presentation explores several use cases for integrating Gemma into applications:

  • Driving the UI with Language: Allowing users to interact with applications using natural language via voice or text, combining traditional UIs with AI. Example: Booking flights by describing travel preferences.
  • Text Generation: Leveraging app context to generate useful content:
    • Data Captioning: Creating AI overviews by processing data within the application and generating summaries. Example: Providing a workout summary in a fitness app.
    • Summarization: Condensing text within the app into a more readable format. Example: Generating social media posts for blog articles.
    • Smart Reply: Generating personalized and nuanced responses in in-app messaging.
    • In-Game Dialogue: Creating custom dialogue for NPCs based on in-game events.

Best Practices for Using Gemma 3 1B

  • Customize the Model: Fine-tune Gemma 3 1B on your own data for better results. A Colab is available to guide this process.
  • Use Few-Shot Prompting: Include examples of inputs and expected outputs in the prompt.
  • Implement RAG (Retrieval Augmented Generation): Use RAG for applications that require querying against large amounts of data.
    • RAG Process:
      1. Chunk the app content (e.g., PDF) into small pieces.
      2. Convert the chunks into embeddings (vector representations).
      3. Store the embeddings in a vector database.
      4. Use a retrieval function to pull out relevant content based on the user's query.
      5. Combine the initial question, retrieved content, and system prompt, and feed it into Gemma.
    • Google AI Edge RAG SDK: An out-of-the-box SDK for on-device RAG on Android (iOS and web coming later).

Implementing Gemma with MediaPipe LLM Inference API

  • MediaPipe LLM Inference API: An easy-to-use API for running Gemma on Android, iOS, and web.
  • Cross-Platform Support: Provides full cross-platform support with optimized performance on both CPU and GPU.
  • Simple Implementation: Requires only a few lines of code to configure, initialize, and run Gemma in your application.
  • Code Examples: Code snippets are provided for Android (Kotlin), iOS (Swift), and web (JavaScript).

Conclusion

The presentation advocates for the strategic use of on-device AI, particularly with models like Gemma 3 1B. It emphasizes the importance of understanding the trade-offs between system GenAI and in-app AI, and provides practical guidance on how to implement and optimize Gemma for various use cases. The key takeaway is that Gemma 3 1B, when properly customized and integrated, can unlock powerful and efficient AI capabilities directly within mobile and web applications, enhancing user experiences while preserving privacy and reducing costs.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.