How a Moonshot Led to Google DeepMind's Veo 3
By Google for Developers
Key Concepts
- VO Project: A Google project focused on generative video models, evolving from Google Brain to DeepMind.
- Generative Models: AI models capable of creating new content, in this case, videos.
- Interleaved Native Audio Capability: A key feature of VO V3, allowing for synchronized audio generation alongside video.
- Moonshot Program: An early initiative within Google Brain focused on ambitious, long-term research goals like video generation.
- Video Prediction Problem: An initial approach to video generation, focusing on predicting future frames based on existing context.
- Text-to-Video Generation: The current primary mode of operation for VO, generating videos from textual prompts.
- Inductive Bias: Assumptions made about data and the world that guide a model's learning process.
- TPUs (Tensor Processing Units): Google's specialized hardware for machine learning computations.
- Human Evaluation: A method for assessing model performance based on human judgment and preference.
- Automated Metrics: Quantitative measures used to evaluate model outputs, often for initial filtering.
- Physics Benchmarks: Using physical plausibility as a metric for video generation quality.
- World Models: AI models that aim to simulate and understand the real world, often for robotics and agent training.
- Pixel vs. Concepts: A debate in AI about whether models should generate raw pixels or higher-level conceptual representations.
- Steerability: The ability of users to control and guide the generation process of AI models.
- Capability Transfer: The process of applying knowledge or advancements from one AI domain (e.g., video understanding) to another (e.g., video generation).
- Gemini Video Understanding: Google's AI model used for analyzing and annotating video content.
Historical Perspective and Project Genesis
The VO project's origins trace back to 2018 within the Google Brain team, initially part of a moonshot program called Brain Video. The core mission was video generation. At that time, generative video models were not widely considered transformative, unlike the current excitement surrounding VO. The team was driven by a desire to push the boundaries of machine learning, tackling challenging problems. Jeff Dean and others at Google supported this endeavor, providing access to significant data and compute resources.
Early Approaches and Challenges
Initial hunches for video generation focused on a video prediction problem. This was considered a more well-posed problem than direct text-to-video generation, as predicting the next frame based on existing context was seen as relatively easier than generating a complex scene from scratch without context. This approach was also rooted in potential applications in robotics, where predicting future environmental states after an action is crucial.
Despite significant progress since 2018, fundamental challenges persist. A major hurdle is the difficulty in evaluating these models effectively. While automated metrics exist to identify completely failing models, hill-climbing metrics (metrics that can be directly optimized) are scarce. Human evaluation is heavily utilized, but it's prone to subjective biases, where factors like increased contrast can lead to better scores without necessarily improving the model's underlying quality. This mirrors issues in LLM evaluation, where longer text can sometimes be favored.
A key difference highlighted between video and text models is the lack of inherent tokenization in video. In text, the world is already tokenized, making it easier to process. In video, the model must first figure out what is important and extract that information to make predictions.
The Evolution of VO: V1, V2, and V3
The VO project was officially launched at Google IO 2024, with V2 released in December 2024 and V3 in May 2024. The journey to these releases involved significant "meandering" and a realization that the required compute power for high-quality video generation took time to catch up with ambitions.
- V1: Described as a high-quality, state-of-the-art model, but it took considerable time to make it scalable for inference and put into users' hands.
- V2: The focus shifted to creating a super high-quality model that could be immediately accessible to users. This marked a significant shift in the team's strategy, moving away from solely publishing research papers ("shipping blog posts") to prioritizing user accessibility. V2 was state-of-the-art for its time and began showing early signs of the audio capabilities that would become a hallmark of V3.
- V3: Launched at Google IO 2024, V3's standout feature is its interleaved native audio capability, generating synchronized audio alongside video. This feature has been a major factor in its widespread acclaim.
The Audio Integration Story
The integration of audio was a long-considered goal. While a paper like VideoPoet (from colleagues) explored video-to-audio and audio-to-video, it didn't focus on joint audio-video generation. The VO team also aimed for this, but in V2, the audio quality did not meet the desired bar. A risky decision was made to delay the audio feature to ensure a good user experience rather than being first to market with subpar audio.
The success of V3's audio, particularly the speech lip-syncing, is credited with creating viral moments, such as the "Yeti blog" and other humorous content. The ability to generate synchronized human-like speech in videos has been a significant differentiator.
User Feedback and Unexpected Trends
The team's excitement during V3's development was tempered by the difficulty of predicting viral success. During internal testing ("dogfooding"), the team anticipated rap battles (like those between Einstein and Newton) to be the viral hit, but the actual popular content revolved around Yeti-related videos and similar unexpected trends.
This experience has reinforced the importance of user feedback. The team has learned that users often desire more control and iteration over generated outputs. This has influenced the roadmap towards enabling greater user control, though balancing this with simplicity and avoiding an overwhelming number of options is a key product challenge.
Technical Aspects and Future Directions
Prompting Paradigms
- Text-to-Video vs. Image-to-Video: While computationally similar, image-to-video generation presents unique learning challenges. Users often want to animate an image in a specific way that might not align with the image's original context, leading to a mismatch between user intent and model capabilities. The concept of "reference-to-video" is explored, where users want to see themselves in a different scene rather than literally animating the starting image.
- Drawing on Images: This technique, discovered internally, is seen as a natural and effective prompting method for image-based video generation, offering a more intuitive level of control than complex text prompts.
- JSON Prompting: The model is not explicitly trained for JSON prompting, and its effectiveness is considered accidental if it works.
Generation Length and Trade-offs
Currently, VO generates 8-second videos. While users desire longer durations, there are significant trade-offs:
- Compute Cost: Longer single-shot generation is more expensive for both inference and training.
- Quality vs. Length: Generating extremely long videos (e.g., an hour) could lead to a convergence to "weird slop" and a loss of coherence, similar to the challenges of long context in LLMs. Maintaining consistency, plot, and character integrity over extended durations is a significant research question.
- User Intent: Many prompts are short, and filling 8 seconds meaningfully can already be challenging. Generating excessively long videos without clear user intent could lead to wasted compute cycles on content that is discarded.
- Extending Videos: The ability to stitch together or autoregressively extend videos is becoming possible, but the quality and coherence of very long generated sequences remain an area of research.
World Models and Physics
The concept of world models, where AI agents explore and learn in virtual environments, is discussed in relation to projects like GD3. For training virtual robots, realistic physics are crucial, unlike the fantastical scenarios often generated by VO. The unresolved research question is whether models should predict literal pixels of the future or a representation of that future, as there's no ground truth for representations.
Capability Transfer and Data Annotation
- Video Understanding to Video Generation: Advancements in Gemini's video understanding capabilities (e.g., V2.5) directly impact VO. Gemini is used to annotate training data with detailed captions, effectively learning the inverse mapping for text-to-video generation. This allows for more verbose and precise descriptions than humans might provide, catering to users who desire fine-grained control over their generated scenes.
- Image Data for Video Models: Surprisingly, incorporating image data alongside video data significantly improves video model training. Images provide a greater diversity of concepts than typical video datasets, helping models learn specific objects and ideas more effectively.
Steerability and Iteration
A key area for future development is natural and economical steerability. The goal is to allow users to easily iterate on generated content without requiring massive compute. Providing a quick rough draft is a desired feature, as users currently have to commit to the full generation cost before knowing if it meets their expectations. The challenge lies in extracting user intent and looping back effectively, especially for sequential events that cannot be captured in short clips.
Future Outlook and User Delight
The VO team aims to continuously improve quality and remain unambiguously state-of-the-art in joint audio-video generation. They are actively seeking feedback from users across various products and cloud services to identify areas for improvement, such as length and steerability. Beyond addressing explicit user requests, the team strives to delight users with novel features they may not yet know they want, leading to the "How did we live without this?" sentiment.
Notable Quotes
- "I don't think anyone in 2018 in the AI world was thinking about generative models of videos as something that would be somehow transformative." - Dumi Erhan
- "We have also not found like a a good way to like the the the the sort of transformer architecture or the inductive bias that is almost certainly correct." - Dumi Erhan
- "We try not to ship blog posts." - Dumi Erhan (referring to the shift towards user-accessible products)
- "If I see a like AI generated video on the internet and has it's it's it's has no sound, I like what's the point? Like I want like to me it's a defect." - Dumi Erhan (highlighting the importance of audio)
- "And then once we ship them, they'll be like, 'How did we live without this until now?'" - Dumi Erhan (on delighting users with unexpected features)
Conclusion
The VO project represents a significant leap in generative video technology, evolving from early research in Google Brain to a state-of-the-art product with integrated audio capabilities. While facing persistent challenges in evaluation and control, the team's focus on user accessibility, iterative improvement, and the integration of novel features like native audio has led to widespread acclaim. Future directions involve enhancing steerability, exploring longer generation lengths, and leveraging advancements in video understanding to further push the boundaries of what's possible in AI-driven video creation.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development