Practical built-in AI with Gemini Nano in Chrome

Chrome for DevelopersAbout 7 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Built-in AI, Gemini Nano, Large Language Model (LLM), Chrome AI, Web AI, Origin Trials, Free Form Prompt API, Writing Assistance APIs (Summarizer, Writer, Rewriter), Translation APIs (Language Detector, Translator), Hybrid SDK, Proofreader API, Multimodal APIs (Image Input, Audio Input), Web Machine Learning Community Group, Hardware Acceleration, Privacy, Interoperability.

Built-in AI in Chrome: An Overview

Chrome's mission is to make Chrome and the web smarter for all developers and users by integrating AI directly into the browser as "built-in AI." This means AI is easily accessible to developers of all skill levels and democratizes access for all users, regardless of device capabilities. By 2025, AI will be ubiquitous, including within Chrome itself.

Built-in AI involves integrating smaller, expert models (like those for translation) and a general-purpose Large Language Model (LLM) called Gemini Nano directly into Chrome. Gemini Nano is downloaded dynamically when needed, rather than being included by default due to its size. These models, along with task-specific fine-tunings, are exposed through high-level APIs, allowing Chrome to match developer intent with the best model for the task. This approach complements lower-level AI model execution methods.

A key advantage of built-in AI is enhanced privacy, as sensitive data remains within the user's browser. It also enables offline functionality and offers potentially superior performance due to hardware acceleration.

Built-in AI APIs: Functionality and Usage

The presentation details several built-in AI APIs, emphasizing that they are in limited availability or experimental stages, often requiring flags to be enabled in Chrome.

1. Free Form Prompt API

  • Description: An origin trial for Chrome extensions and testable behind a flag for the web, allowing developers to experiment with LLMs.
  • Usage:
    • Create a session with the LLM using languageModel.create.
    • Call the prompt function on the session object, passing a prompt (e.g., "Explain quantum computing like I'm 5").
    • In Chrome extensions, use chrome.ai.originTrial.languageModel.create to create the session.
  • Example: Creating a Chrome extension that explains website content like "Explain like I'm 5".

2. Writing Assistance APIs

  • Family: Includes the Summarizer API (origin trial), Writer API (behind a flag), and Rewriter API (behind a flag).
  • Summarizer API:
    • Usage:
      • Get the text to be summarized.
      • Create a summarizer instance using summarizer.create, specifying the type (e.g., "headline") and desired length (e.g., "short").
      • Call the summarize function to get a headline suggestion.
  • Writer API:
    • Usage:
      • Provide a shared context (e.g., "things that got better").
      • Create a writer instance using writer.create, passing the shared context.
      • Call the write function, specifying the desired output (e.g., "write a blog post about the world changing for the better").
  • Rewriter API:
    • Usage:
      • Provide an existing draft and a shared context.
      • Create a rewriter instance using rewriter.create, passing the shared context.
      • Call the rewrite function, specifying the desired tone (e.g., "less fun").
  • Example: A page that lets users write blog posts using the summarizer API to suggest headlines, the writer API to draft the blog post, and the rewriter API to rewrite the draft.

3. Translation APIs

  • Family: Includes the Language Detector API (origin trial) and the Translator API (origin trial).
  • Language Detector API:
    • Usage:
      • Create a language detector using create.
      • Call the detect function, which returns an array of language detection objects with confidence scores and detected languages.
  • Translator API:
    • Usage:
      • Create a translator using create, specifying the source and target languages.
      • Call the translate method to get the translation.
  • Example: An insurance assistant chatbot that uses the language detection API to identify the user's input language and the translator API to translate the input to English and the model's response back to the user's native language.

4. API Consistency and Model Availability

  • All APIs follow a consistent creation pattern using a create function, often with optional options.
  • The availability function provides a consistent way to check the readiness of the model:
    • unavailable: Implementation doesn't support the task/options.
    • downloadable: Implementation supports the task/options, and the model can be downloaded.
    • downloading: Download is in progress.
    • available: API is ready to use.

5. Proofreader API

  • Description: A new API based on fine-tuning the Gemini model for proofreading tasks.
  • Functionality: Detects and corrects spelling issues, punctuation errors, capitalization errors, wrong prepositions, skipped words, and general grammar issues.
  • Usage:
    • Instantiate the proofreader using the create function, optionally specifying correction types and explanations.
    • Call the proofread function, which returns an array of detailed correction objects and a complete corrected version of the input string.
  • Availability: Testable behind a flag in Chrome 139 or 140.

6. Multimodal APIs (Image and Audio Input)

  • Image Input:
    • Allows the Gemini model to analyze visual content.
    • Use Cases: Generating alternative text for images, creating product description drafts from product images, extracting textual information from images (OCR).
    • Usage:
      • Create a session using create, passing type: 'image' in the expectedInputs array.
      • Pass an image element, blob, bitmap, video still frame, or canvas object as input.
      • Prompt the API about the inputs.
  • Audio Input:
    • Allows the Gemini model to analyze audio content.
    • Use Cases: Creating transcription suggestions, adding textual search for audio segments, classifying audio files.
    • Usage:
      • Create a session using create, passing type: 'audio' in the expectedInputs array.
      • Pass a blob of an audio file or a reference to an audio element as input.
      • Prompt the API to transcribe the speech.
  • Availability: Image and audio input for the Prompt API will be entering origin trial in either Chrome 139 or Chrome 140.

Hybrid SDK

  • Challenge: AI model execution is demanding, and there's a support gap across platforms (Chrome OS, Android, desktop devices, other browsers).
  • Solution: A hybrid SDK, extending the Firebase web SDK, will use built-in APIs when available and fall back to Gemini on the server.
  • Availability: A developer preview will be available in Chrome, with the full solution coming later this year.

Partner Use Cases

The presentation highlights various partners who have tested and implemented the built-in AI APIs:

  • Redbus: Summarizes bus reviews using the Summarizer API.
  • Moravia: Summarizes product reviews using the Summarizer API.
  • Brightides: Integrates the Summarizer API into its content management system for personalized news summaries.
  • Terra: Combines the Summarizer API and Translator API to summarize and translate news stories.
  • Cyber Agent: Developed a Chrome extension for Amoeba Blog, providing AI assistance for generating titles, headings, and paragraphs.
  • Policy bazar: Built an insurance assistant chatbot that uses the Language Detection API and Translator API to support Indic languages.
  • Make My Trip: Uses the Writer API to expand customer reviews based on keywords and the Rewriter API to refine them.
  • Deote: Experimenting with the Prompt API, Summarizer API, Writer API, and Rewriter API for form filling, engineering onboarding, and feedback.
  • Nykaa: Experimented with visual search based on the Prompt API with image input.
  • Chefkoch: Used the Prompt API with image input to digitize handwritten recipes.
  • The Independent: Integrated the Prompt API with image input into their content management system to generate image titles and descriptions.
  • Grupo OX: Used the Prompt API with audio input to transcribe audio messages in their chat platform.

Community Engagement

The Chrome team has actively engaged with the developer community through an early preview program with over 16,000 members and a built-in AI challenge hackathon with nearly 9,000 participants.

Web Standards and Interoperability

The Chrome team is working with the Web Machine Learning Community Group at the W3C to standardize the APIs and ensure interoperability across browsers.

Conclusion

Chrome is actively integrating AI into the browser through built-in APIs, offering developers powerful tools for various use cases. The APIs are designed with consistency and privacy in mind, and the Chrome team is committed to community engagement and standardization. The upcoming Hybrid SDK will further extend the reach of these APIs. Developers are encouraged to explore the APIs, provide feedback, and contribute to the standardization effort.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.