Release Notes: Gemini's multimodality

Google for DevelopersAbout 7 min readJul 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Multimodal AI
  • AGI (Artificial General Intelligence)
  • Vision AI
  • Native Multimodality
  • Token Representation
  • Video Understanding
  • OCR (Optical Character Recognition)
  • Spatial Understanding
  • Temporal Reasoning
  • Model Capability Transfer
  • Visual Reasoning
  • Embodied Reasoning
  • Layout Preserving Transcription
  • Model Behavior
  • Proactivity
  • Bidirectional Audio/Video Interface

Gemini: Multimodal Vision Capabilities and Future Directions

Multimodal AI and the Importance of Vision

  • Gemini was designed from the outset as a multimodal model, recognizing that Vision is a core component of human experience and essential for achieving AGI.
  • Many tasks across domains like medicine and finance have a strong visual component.
  • The goal is to create models that can see and perceive the world like humans, enabling them to perform tasks in a more natural and effective manner.

Native Multimodality: Training and Representation

  • Gemini is a natively multimodal model, meaning it's trained from the ground up on multiple modalities simultaneously.
  • Text, images, video, and audio are converted into a token representation, and the model is trained on all this information together.
  • This allows the model to understand not just text, but text in conjunction with images, audio, and video.
  • The abstraction is that these models should be able to see and perceive the world like we do.

Challenges and Trade-offs in Multimodal Training

  • Information loss is a significant research problem. Converting images into token representations inherently loses some information.
  • Video sampling (e.g., one frame per second) also results in information loss, as the model doesn't see the entire video stream.
  • Despite these losses, models generalize surprisingly well after seeing enough images and videos.
  • Research is ongoing to develop less lossy image representations.

Gemini 2.5 Pro: Advancements in Video Understanding

  • Gemini 2.5 Pro demonstrates state-of-the-art model performance in video understanding.
  • Previous Gemini models had robustness issues with long videos, tending to focus on the initial segments.
  • Improvements in core Vision capabilities also contribute to better video understanding.
  • A key capability is video to code conversion, enabling the creation of animations, websites, and interactive learning applications from videos.
  • Example: Converting a YouTube recipe video into a step-by-step recipe or a lecture video into lecture notes and a web page.

Interplay of Vision Capabilities: Capability Transfer

  • A key advantage of a single multimodal model is positive capability transfer across different Vision tasks.
  • Improvements in one area (e.g., code generation) benefit other areas (e.g., video to code).
  • Bundling capabilities like OCR, detection, and segmentation into Gemini eliminates the need for separate models.
  • Example: Transcribing a video requires strong OCR and temporal understanding.
  • Peer programming use case: Streaming a video of an IDE to Gemini for code-related assistance.

Prioritizing Model Development: Use Cases and AGI

  • Model development is prioritized based on:
    • Critical use cases for current users and customers (e.g., API users, Google products).
    • Long-term aspirational capabilities essential for building powerful AI systems and AGI (e.g., visual reasoning).
    • Unexpected capabilities that emerge from scaling models (e.g., image to code, video to code).
  • Visual reasoning example: Gemini analyzing the trajectory of a ball in a pinball machine image. This capability is seen as crucial for robotics and self-driving cars.
  • UX to code example: Generating a prototype from a UX sketch using HTML, JavaScript, or React.

Vision Use Cases: Beyond Human Capabilities

  • Vision use cases are categorized into three buckets:
    • Tasks that existing models/systems could do (e.g., OCR, translation, image retrieval).
    • Tasks that a human expert could do (e.g., answering questions about surroundings using Gemini while walking around a city).
    • Tasks beyond human capabilities or feasible timeframes (e.g., generating a highlights reel from a six-hour sports game, video to interactive learning application).

The "Everything is Vision" Mantra: Product Development

  • The future of AI products involves embracing the "Everything is Vision" concept.
  • This means building products that can see the world like humans and act as domain experts in every field.
  • Anthropomorphizing models as expert humans can guide interface design.
  • Moving beyond turn-based interactions to bidirectional audio/video interfaces for more natural communication.
  • Proactivity: AI systems that can anticipate needs based on visual cues and suggest actions.
  • Example: An AI system that monitors a user's screen and proactively offers solutions when an error occurs in the terminal.
  • Multitasking: AI systems that can process audio and video while simultaneously thinking and performing actions on a screen.
  • Example: Gemini looking at what you're doing as you're cooking and then proactively, based on visual cues in the video, suggests for things to do, right? So I dunno, I was boiling pasta, it was like, hey, add the pasta now and things like that.

Video Understanding: Technical Details

  • Gemini is a foundation model capable of state-of-the-art video understanding and reasoning.
  • It understands both the audio and visual components of a video.
  • Audio and frames corresponding to that audio are interleaved at each time chunk.
  • This approach generalizes well, allowing the model to understand videos effectively.

Frames Per Second (FPS) and Tokenization

  • The model was initially designed and trained at one FPS, which worked well.
  • Higher FPS is beneficial for use cases like analyzing golf swings or dance moves.
  • Customers have been slowing down videos to achieve higher FPS for analysis.
  • The initial one FPS design was partly due to tokenization limitations, allowing for approximately one hour of video.
  • More efficient tokenization methods have been developed, enabling up to six hours of video with 2 million contexts.
  • With Gemini 1.0, representing an image with 64 tokens was just a very lossy representation. What we see today is actually 64 tokens performs remarkably well, almost actually to the same quality level as 256.

Future Directions: Output Modalities and Spatial Understanding

  • The goal is to create models that excel at both multimodal input and output.
  • On the Vision side, there's a focus on bringing capabilities together to form a more cohesive system.
  • Spatial understanding (generating 2D/3D bounding boxes, point coordinates, segmentation masks) is a key area.
  • Example: Using Gemini to detect the person furthest to the left in an image or identify the drink with the fewest calories in a fridge.
  • Spatial understanding is a core building block for embodied reasoning and perception in robotics.

Document Understanding and OCR

  • Documents are a powerful medium of information, and Gemini is designed to analyze and reason over them effectively.
  • Gemini combines OCR capabilities with its reasoning backbone.
  • Key use case: Feeding a large number of documents as context to the model to perform complex multi-step tasks.
  • Gemini can see a document like a human, understanding its format, charts, images, and diagrams.
  • Example: Analyzing earnings reports from companies over the last 10 quarters (millions of tokens) to perform company analysis.
  • Layout preserving transcription: Gemini can transcribe a document while preserving its layout, style, and structure.

Vision as a Store of Information

  • Vision unlocks visual information, making it more accessible and useful.
  • Example: Cataloging books in a library by genre and author using a video of the bookshelf.
  • Example: Cataloging snacks in micro kitchens.

Team Collaboration and Future Vision

  • The Gemini multimodal team is a collaborative group focused on bringing capabilities together into a single model.
  • The team emphasizes a close product model feedback loop, incorporating developer and consumer needs into model development.
  • They also focus on building capabilities that will be relevant in the future, anticipating how people will interact with these models in the years to come.

Transition to Model Behavior

  • The focus is shifting towards making models feel more natural to interact with.
  • This involves giving models skills like empathy, understanding user intent, and developing a personality.
  • Exploring visual formats for communicating information in a more information-dense manner.
  • The goal is to create AI systems that are likable and easy to interact with.

Conclusion

The development of Gemini and its multimodal capabilities represents a significant step towards AGI. By focusing on Vision as a core component of intelligence, training models on multiple modalities simultaneously, and prioritizing both current use cases and long-term aspirational goals, the Gemini team is pushing the boundaries of what's possible with AI. The emphasis on collaboration, user feedback, and a future-oriented vision ensures that these models will continue to evolve and become increasingly powerful and useful in a wide range of applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.