AI Stage - Day 2 (Google I/O 2025)

Google for DevelopersAbout 6 min readMay 22, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Gemini Models: Multimodal AI models developed by Google DeepMind, capable of processing text, image, video, audio, and code.
  • Gemini API: A programmatic interface for accessing Gemini models, offering a low barrier to entry for developers.
  • Gemini TTS: A model for generating high-quality audio from text, with customizable voices, languages, and moods.
  • Gemini Live API: A real-time, low-latency API for interactive experiences, supporting native audio input and output.
  • Deep Think: An advanced mode for Gemini 2.5 Pro, allowing the model to consider multiple answers before providing the best one.
  • Gemini Diffusion: A diffusion architecture for faster image generation.
  • AI Studio: A no-code UI for experimenting with Gemini models and developing solutions.
  • Function Calling: The ability for Gemini models to call external tools and functions to enhance their capabilities.
  • Agentic Capabilities: The ability for AI models to act autonomously and take actions to assist users.
  • Thinking Budget: A feature for Gemini 2.5 Flash that allows developers to control the amount of reasoning the model performs.
  • Thought Summaries: A feature that provides insights into the reasoning steps the model took to arrive at a conclusion.
  • URL Context: A tool for extracting in-depth content from webpages.
  • Gemma: An open-source family of models based on Gemini, available in various sizes and flavors.
  • Gemma 3: The latest version of Gemma, featuring multimodal capabilities, multilingual support, and a longer context window.
  • Gemma 3n: A new model that can run on as little as 2 gigabytes of RAM.
  • MedGemma: A Gemma variant for the health industry.
  • SignGemma: A Gemma variant for sign language.
  • ShieldGemma: A model of safety clarify.
  • Quantization: A technique for reducing the size of AI models by reducing the precision of their parameters.

1. Gemini Models and API Overview

  • Gemini's Multimodal Nature: Gemini models are designed from the ground up to handle multimodal data, including text, images, video, audio, and code. This allows developers to create applications that can process and understand information in various formats.
  • Model Families:
    • Gemini 2.5 Pro: The most powerful model, suitable for complex tasks requiring deep reasoning and coding.
    • Gemini 2.5 Flash: Offers the best price-performance ratio.
    • Gemini 2.0 Flash-Lite: A small, fast, and cheap model for high-volume tasks like summarization.
    • Gemini Nano: A smaller version of Gemini designed to run locally on devices like Android phones.
    • Gemini Embedding: Used for creating high-quality embeddings for semantic ranking and organization.
  • API Access and Free Tier: The Gemini API provides access to all public models, including Gemini, Gemma, and others. It offers a generous free tier for experimentation. SDKs are available for Python, JavaScript, and Go.
  • Google AI Studio Integration: Google AI Studio allows developers to test API capabilities before building applications at scale. It now includes co-generation features.

2. New Features and Capabilities

  • Gemini TTS Model: Generates high-quality audio from text with customizable voices, languages, and moods.
  • Native Audio Output Models: Available through the Gemini Live API, these models offer more natural-sounding voices and better contextual understanding. They support seamless language transitions and are available with thinking enabled for complex use cases.
  • Deep Think Mode: An advanced mode for Gemini 2.5 Pro that allows the model to consider multiple answers before providing the best one.
  • Gemini Diffusion: A faster diffusion architecture for image generation.
  • YouTube Link Analysis: The Gemini API can now analyze information from YouTube links.
  • Dynamic Frame Rate Support: The API supports dynamic frame rates per second for video processing.
  • Video Clipping and Image Segmentation: New features for video and image analysis.
  • Long Context and Context Caching: Gemini models support long context windows (up to 2 million tokens). Explicit and implicit context caching are available to reduce input token pricing.
  • Bounding Boxes and Image Segmentation: The API can provide bounding boxes and image segmentation data, including classifications and mask segmentation.
  • Streaming Support: The API supports streaming for real-time applications.
  • Image Generation and Editing: Gemini Image Out allows for generating images from text and editing existing images through chat interactions.
  • Veo 3 Integration: Veo 3, which supports text-to-video and image-to-video generation, will soon be available through the API.
  • Proactive Audio and Effective Dialogue: New features for the Live API that allow the AI to proactively respond and pick up on the user's tone and sentiment.
  • Asynchronous Function Calling: A new feature that allows for running functions in the background while maintaining a real-time conversation.

3. Agentic Capabilities and API

  • Agent Architecture: Agents typically consist of an orchestration layer, a model layer, and a tools layer.
  • Key Components:
    • Planning and Reasoning: Gemini 2.5 series models are trained for planning and reasoning.
    • Tools: Google Search, code execution, URL context, and function calling are available as tools.
    • Orchestration: Involves defining the agent's behavior, goals, memory, and reasoning process.
  • Agentic Primitives: The Gemini API provides high-quality primitives for building agents, including the 2.5 series models, Deep Think, and thinking budgets.
  • Thinking Budget and Thought Summaries: Gemini 2.5 Flash allows developers to set a thinking budget and receive thought summaries, providing insights into the model's reasoning process.
  • Code Examples: The presentation includes code snippets demonstrating how to use the Gemini API with the Python SDK to access thinking config, URL context, and other features.

4. Gemma Open Models

  • Gemma Family: An open family of models based on Gemini, designed for accessibility and collaboration.
  • Gemma 3 Features:
    • Multimodal: Supports image and video input.
    • Multilingual: Supports over 140 languages.
    • Longer Context Window: Up to 128K tokens.
  • Gemma 3n: A new model that can run on as little as 2 gigabytes of RAM.
  • Community Feedback: The development of Gemma is driven by community feedback, with features like larger context windows and smaller models being added in response to user requests.
  • LMArena Benchmark: Gemma models perform well on the LMArena benchmark, with the 27B model being one of the top open models.
  • Quantization: Techniques for reducing the size of Gemma models while preserving quality.
  • Fine-Tuning: Gemma models can be fine-tuned for specific tasks and domains.
  • Community Variants: The community has built over 70,000 Gemma-based models.
  • MedGemma: A Gemma variant for the health industry.
  • SignGemma: A Gemma variant for sign language.
  • ShieldGemma: A model of safety clarify.

5. Building with Gemma

  • Google AI Studio: A platform for experimenting with Gemma models and deploying them to Google Cloud Run.
  • Open Source Tools: Gemma is designed to be easily integrated into open-source tools like Keras, vLLM, and Ollama.
  • Kaggle Integration: Gemma models are available on Kaggle, where users can discover models, share knowledge, and participate in AI challenges.
  • Quantization Techniques: Techniques for reducing the size of AI models by reducing the precision of their parameters.

6. Conclusion

The presentation highlights the new features and capabilities of the Gemini models and API, as well as the Gemma open models. It emphasizes the importance of multimodal understanding, agentic capabilities, and community collaboration. The speakers encourage developers to start building with these tools and provide feedback to help shape the future of AI.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.