Veo 3 for Developers - Paige Bailey

AI EngineerAbout 5 min readJun 22, 2025Watch original
THE SUMMARYAI-generated

Generative Media: V3, Imagine 4, and LIA 2 - Google DeepMind

Key Concepts:

  • Generative Media: V3 (Video & Audio), Imagine 4 (Images), LIA 2 (Music)
  • Prompt Adherence
  • Native Audio Generation
  • Stylistic Consistency
  • Contextual Consistency
  • Reference Powered Videos
  • Outpainting
  • Character Control
  • Interpolation (First & Last Frame)
  • SynthID Watermarks
  • Granular Creative Controls
  • Modality
  • Tokens

V3: Video and Audio Generation

  • Overview: V3 is Google DeepMind's new state-of-the-art video generation model that natively combines video and audio generation. This means the model composes tokens across multiple modalities, unlike previous models where audio was added as a separate tool.
  • Key Features:
    • Native Audio Generation: V3 generates audio (music, sound effects, background noises) directly within the model, not as a post-processing step.
    • Improved Consistency: V3 demonstrates significant improvements in stylistic and contextual consistency compared to previous models. This addresses issues like characters inconsistently appearing across frames or backgrounds changing unexpectedly.
    • Prompt Adherence: V3 is better at capturing the nuances of input prompts, resulting in videos that more closely match the desired outcome.
  • Underlying Research: V3 is built upon years of research, including models like GQN and Walt.
  • Responsibility: V3 incorporates responsibility measures, including human-visible watermarks and SynthID watermarks for synthetically generated content.
  • Access:
    • Google AI Ultra plan
    • Google AI Pro subscribers (limited uses via Gemini mobile app)
    • Private preview via Vertex AI
    • Early access form (QR code provided)
  • Code Sample: The presentation included a code sample demonstrating the ease of use of the V3 API, highlighting how a few lines of code can specify the output bucket, input image, and aspect ratio.

V2 Enhancements

  • Creative Control: Recent updates to V2 focused on enhancing creative control, including:
    • Reference Powered Videos: Allows users to combine a person or environment from one video with another, creating stylized compositions.
      • Example: Inserting a monster character into various environments based on text descriptions.
      • Benchmarks show V2 performing well compared to Runway Gen 4 and Kling in reference-powered video generation.
    • Style Matching: Users can upload a reference image to influence the style of the generated video.
    • Camera Controls: Natural language commands can control camera movements like zoom, pan, and rotation.
    • Outpainting: Extends existing video frames to imagine the surrounding scene.
      • Example: Used in a project with the Sphere around Wizard of Oz.
    • Object Manipulation: Ability to add or remove objects from scenes.
    • Character Control: Control character movements, lip-syncing, and voice tone.
      • Users can add a script and voice tone to have the character produce sound consistently with the location.
    • First and Last Frame Interpolation: V2 can interpolate between a starting and ending image to create a video sequence.

Imagine 4: Image Generation

  • Overview: Imagine 4 is Google DeepMind's image generation model.
  • Key Features:
    • Realism: Preserves realism in generated images, including humans, animals (e.g., puppies), and objects.
    • Detail Preservation: Maintains detail across generated images.
    • Diverse Styles: Supports a wide range of styles.
    • Typography: Capable of generating images with typography.
  • Artist Collaboration: The team has been collaborating with artists like Ross Lovegrove.

LIA 2: Music Generation

  • Overview: LIA 2 is Google DeepMind's high-fidelity music and professional-grade audio generation model.
  • Key Features:
    • Granular Creative Controls: Provides granular creative controls to steer inputs, outputs, tones, and styles of music.
    • Music AI Sandbox: A visual interface similar to Ableton for music creation.
    • Music Effects: A project from the labs team that allows users to compose beats using natural language.
  • Real-time Collaboration: LIA 2 has been developed in collaboration with musicians like Jacob Collier and Toro y Moi.
  • Integration with Education: Jacob Collier's quote highlights the potential of AI music tools to democratize music education by providing access to the "whole of music" from day one.
  • Digital Watermarking: Incorporates techniques like SynthID for digital watermarking of generated assets.

Evolution of Generative Media

  • Comparison of Text-to-Video Models (2022-2024): The presentation showcased the rapid progress in text-to-video generation by comparing outputs from different models over the past few years, using the prompt "a raccoon wearing a black jacket dancing in slow motion in front of the pyramids."
    • Walt (2023): Choppy and short video.
    • Cling 2.0 (2024): Improved but still limited.
    • V2 (2024): Cute raccoon, but limited dancing ability.
    • V3: Stylish raccoon with detailed environment and realistic movement.
  • Image-to-Video Transformation: V3 can transform static images into dynamic video content, steered via natural language prompts.

V3 Demonstration: Replicating a Commercial

  • Original Commercial: A Chick-fil-A commercial featuring a person named Paige describing what makes the chicken sandwich original.
  • V2 Replication Process:
    1. Use Gemini to create a detailed plan from the original video.
    2. Segment the video into prompts to handle the 8-second limitation.
    3. Use Video Music Effects to create a background track.
    4. Stitch everything together in Camtasia with transitions.
  • V3 Replication Process:
    1. Generate a description of the original video.
    2. Give the description to V3.
  • Result: V3 was able to replicate the commercial with a single prompt, demonstrating its ability to understand and generate complex video content.

Conclusion

V3 represents a significant leap forward in generative media, particularly in video and audio generation. Its ability to natively combine modalities, maintain consistency, and adhere to prompts opens up new possibilities for creative expression and content creation. Google DeepMind is committed to expanding access to V3 and adding controls around durability. The presentation also highlighted the importance of collaboration with artists and musicians in developing these technologies responsibly.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.