Veo 3.1 fully tested
By AI Search
Key Concepts
- VO 3.1: Google's latest AI video generator.
- Ingredients to Video: A feature allowing users to upload reference images (characters, objects) to be incorporated into generated videos.
- Extend Feature: A method to create longer videos by stitching together multiple 8-second clips, using the last frame of one as the first frame of the next.
- Flow: Google's native platform for video generation using VO models.
- V3.1 Fast vs. V3.1 Quality: Two modes within Flow, offering different generation speeds and quality levels.
- Prompt Adherence: The AI's ability to accurately follow instructions given in text prompts.
- Character Consistency: The AI's capability to maintain the appearance of characters across different generated scenes.
- Sora 2: A competing AI video generation model used for comparison.
- Cling 2.5 & High Law O2: Other AI video generation models mentioned for their performance in specific areas.
- Hugging Face, Foul.ai, Waves, Replicate: Platforms where VO 3.1 and other AI models can be accessed.
Google VO 3.1: A Detailed Review and Comparison
This video provides an in-depth test and comparison of Google's new AI video generator, VO 3.1, highlighting its capabilities, limitations, and how it stacks up against other leading models.
Official Announcement and Core Claims
Google's official announcement for VO 3.1 emphasizes incremental improvements over its predecessor, V3. Key claims include:
- Richer Audio: Enhanced sound generation capabilities.
- More Narrative Control: Improved ability to guide the story and sequence of events.
- Enhanced Realism: Higher fidelity and more lifelike visuals.
- Stronger Prompt Adherence: Better understanding and execution of user prompts.
- Improved Quality: Overall enhancement in the visual output.
Key Features and Limitations
1. Video Length:
- Native Generation Limit: VO 3.1 is currently limited to generating videos of 8 seconds in length.
- Extend Feature: While not native, longer videos can be created by using the "extend" feature. This process involves taking the last frame of an 8-second clip and using it as the first frame for the next generated 8-second clip, effectively stitching them together. This is described as "not ideal" and not a native solution for longer video generation.
2. Ingredients to Video Feature:
- This is highlighted as one of VO's most powerful features. It allows users to upload multiple reference images (e.g., characters, objects, backgrounds) to be incorporated into the generated video.
- Example 1 (Destiny Gundam): The presenter uploaded an image of the Destiny Gundam and a desert background. The prompt was to "zoom into this character running in this desert scene. cinematic fast camera movement." The result showed the Gundam in the desert, with the presenter noting that while not perfect, the character and background were generated well, even with the Gundam's complex design.
- Platform Specifics (Flow vs. Higsfield):
- Google Flow: When using images as ingredients in Flow, users are forced to crop them to either landscape or portrait aspect ratios, which can be problematic for square images.
- Higsfield: This platform allows uploading multiple images without mandatory cropping, making it more flexible for certain use cases.
- Example 2 (Influencer Product Video): Using Higsfield, three product images (headphones, handbag, sneakers) were uploaded. The prompt requested a "Tik Tok style influencer video. She is holding up and talking about these three products one by one. Low quality amateur video taken on a phone." The generated video successfully showed an influencer presenting each product sequentially, demonstrating its utility for product promotion.
- Example 3 (Anime Character Consistency): Two photos of anime characters with complex outfits were uploaded. The prompt was to have them "dance in sync on stage. It's going to be fast movements and upbeat." The generation showed good character consistency in appearance and outfits, though some minor distortions and warping were observed near the edges.
Performance Testing and Comparisons
The video then moves into testing VO 3.1 with more challenging prompts and comparing its output with Sora 2, Cling 2.5, and High Law O2.
1. Emotional Expression Changes:
- Prompt: "A woman to laugh really hard and then she looks shocked. Then she bursts out crying, then she looks really excited."
- VO 3.1 Result: The video showed realistic transitions through these emotions within the 8-second limit.
- Sora 2 Comparison: Both models performed well, with the presenter asking viewers to choose their preference.
2. World Understanding and Character Generation:
- Prompt 1 (Pikachu ASMR): "Closeup of Pikachu doing ASMR."
- VO 3.1 Result: Generated a more realistic, less recognizable version of Pikachu. The presenter found it "not bad" but not ideal.
- Sora 2 Result: Generated a more iconic Pikachu, with the presenter preferring Sora 2's output.
- Prompt 2 (Lord of the Rings, Gen Z Style): "Lord of the Rings, but Gen Z style."
- VO 3.1 Result: Did not generate recognizable Lord of the Rings characters and was deemed "definitely not as good as Sora 2."
- Sora 2 Result: Generated a video with Gen Z slang and themes applied to the Lord of the Rings concept, which was significantly better.
- Prompt 3 (Goku vs. Mewtwo): "Let's get Goku to fight Mewtwo in a crowded stadium."
- VO 3.1 Result: Goku looked "okay" but had noise and warping. Mewtwo was "completely wrong."
- Sora 2 Result: Generated accurate representations of Goku and Mewtwo in a battle scene, significantly outperforming VO 3.1.
- Prompt 4 (Anime Characters, Chinese New Year): "Chinese New Year performance with Naruto, Gojo, Nezuko, and One Punch Man."
- VO 3.1 Result: Generated a realistic video, correctly identifying Naruto and partially Gojo, but failing to generate Nezuko and One Punch Man. It also exhibited errors like two Narutos.
- Sora 2 Result: Generated a more cohesive and accurate representation of the requested characters.
- Prompt 5 (Japanese Commercial): "A Japanese commercial of a food delivery service where your food is delivered by old ugly women in maid costume."
- VO 3.1 Result: Did not produce a commercial and failed to meet the prompt's specific requirements.
- Sora 2 Result: Generated a good commercial that aligned with the prompt.
- Prompt 6 (Comedian's Joke): "A comedian tells a joke on stage, but it wasn't funny and there was awkward silence from the crowd."
- VO 3.1 Result: Failed to follow the prompt; the crowd was laughing instead of silent.
3. Gameplay Scene Generation:
- Prompt (Starcraft): "Starcraft gameplay of Teal Protoss versus Red Zerg."
- VO 3.1 Result: Did not resemble a Starcraft scene and had incorrect details.
- Sora 2 Result: Provided a more accurate and detailed Starcraft scene.
4. Physics and Anatomy Understanding:
- Prompt 1 (Unicyclist Juggling): "A man riding a unicycle and juggling red balls."
- VO 3.1 Result: Juggling was "absolutely awful."
- Cling 2.5 & High Law O2 Comparison: These models were cited as being "way cheaper and faster" and handling physics "a lot better."
- Prompt 2 (Gymnast Flip): "A gymnast doing a flip on a balance beam."
- VO 3.1 Result: Better than V3, with limbs remaining intact, but the action was "still kind of weird" and not a proper flip.
- Cling 2.5 Comparison: Cling 2.5's generation looked "a lot better." The presenter recommends Cling 2.5 or High Law for anatomically tricky scenes.
- Prompt 3 (Breakdancing): Testing VO 3.1's ability to handle breakdancing.
- VO 3.1 Result: "Just not really good at handling these types of prompts."
- Cling 2.5 Comparison: Cling 2.5 was significantly better and also cheaper/faster.
- Prompt 4 (3D Animation - Princess and Dragon): "A princess wearing a glittery white dress running away from a massive red dragon with glowing red eyes. 3D Disney Pixar style."
- VO 3.1 Result: "Not bad," followed the prompt mostly, but movements were "very slow."
- Cling & High Law Comparison: Both looked "a lot better."
- Prompt 5 (Complex Scene with Multiple Elements): "A ballerina spinning in a studio with mirrored walls with scattered point shoes and sheet music. A rabbit watches a top a grand piano. Outside the large window, an elephant balances on a circus ball."
- VO 3.1 Result: Significant errors: two rabbits, messed-up ballerina face, elephant not balancing, elephant not outside the window.
- High Law O2 Comparison: High Law's generation was "astounding" and significantly better. The presenter reiterates that High Law or Cling are better for tricky physics or anatomy.
5. Text Generation and Diagrams:
- Prompt 1 (Camera Movement and Text Overlay): "A camera pushing in towards a couple kissing and then it tilts up to show the sky and then it should overlay this text."
- VO 3.1 Result: Not a continuous scene, faded to the sky, and the text was incorrect. The presenter states VO 3.1 "doesn't seem to generate text very well."
- Prompt 2 (Pythagorean Theorem Diagram): "A professor explaining the Pythagorean theorem on the whiteboard. As you can see here, when we square sides A and B, they perfectly equal the square of the hypotenuse C."
- VO 3.1 Result: Did not understand the diagram. The presenter notes that other leading AI models also struggle with this.
6. Prominent People and Censorship:
- Prompt (Will Smith Eating Spaghetti): "Will Smith eating spaghetti. Absolutely delicious. I haven't had spaghetti this good in a long time."
- VO 3.1 Result (Direct Prompt): Due to Google's "pretty strict censorship," it was unable to generate the desired Will Smith.
- VO 3.1 Result (Frames to Video): Using the "Frames to Video" feature (similar to "Ingredients to Video"), the presenter attempted to upload a photo of Will Smith. However, Google Flow policies prohibit uploading "prominent people at this time."
- Higsfield Result: Higsfield successfully allowed the upload of the Will Smith image. The prompt "M, I can't get enough of this. And eats the spaghetti" resulted in a good generation where he eats spaghetti realistically and speaks the specified dialogue.
- Sora 2 Comparison: Sora 2 does not allow uploading photorealistic people's images.
7. Image to Video with Audio:
- Prompt (K-Pop Group): Uploaded an image and prompted for "a K-pop group singing and dancing on stage. It's going to be an upbeat Korean pop song."
- VO 3.1 Result: Did not generate a Korean song but the dancers moved in sync. Movements were "not too bad," but there was noticeable warping and deformation.
- Prompt (Warrior vs. Monster): Uploaded an image and prompted for a warrior sprinting, leaping, and being engulfed in flames by a monster, with "epic orchestral battle music with epic sound effects."
- VO 3.1 Result: Movements were "really slow" and not good for high-action scenes.
- Higsfield (High Law O2) Comparison: High Law looked "much better" for epic cinematic action scenes, especially if audio is not a primary concern.
8. Consistency with Complex Images:
- Prompt: Uploaded a complicated image with many people and details, leaving the prompt empty to test consistency.
- VO 3.1 Result: "Not great." People appeared and disappeared, and there was considerable warping on faces.
9. Anime Dialogue Generation:
- Prompt: Uploaded an anime image and specified dialogue for the characters.
- VO 3.1 Result: The generated video appeared to have the characters speaking the specified Japanese dialogue, with fluid and natural movements. The presenter noted understanding "kimo."
- Sora 2 Comparison: Sora 2 would not even allow this generation.
10. Motion Graphics and Voiceover:
- Prompt: "A motion graphic highlighting the country of Estonia. I also want a voice over of a woman with an Indian accent describing the history of Estonia."
- VO 3.1 Result: Failed on multiple fronts:
- Incorrect location for Estonia.
- Voiceover was not an Indian accent.
- Text was incorrect.
- Comparison: The presenter notes that none of the other leading video models could get this prompt correct either.
- VO 3.1 Result: Failed on multiple fronts:
Creating Longer Videos with the Extend Feature
- The "extend" feature is the only method for creating videos longer than 8 seconds in VO 3.1.
- Process in Google Flow:
- Generate an 8-second video.
- Hover over the video and click "Add to Scene."
- In the scene builder, click the "+" sign.
- Select "Extend."
- Prompt further for the next 8-second segment.
- Mechanism: The feature uses the last frame of the previous video as the first frame of the new one.
- Limitation: This does not create a seamless long video and VO 3.1 does not natively generate videos longer than 8 seconds.
Synthesis and Conclusion
- Overall Assessment: VO 3.1 is a "slight incremental upgrade" from V3.
- Strengths:
- Handles character consistency very well, making the "Ingredients to Video" feature powerful for incorporating specific objects or characters.
- Generates video with audio.
- Weaknesses:
- Fails significantly in areas requiring precise physics or anatomy (breakdancing, gymnastics).
- Cannot generate text or diagrams effectively.
- World understanding and generation of existing characters are not as good as Sora 2.
- Limited to 8-second native generation.
- Performance issues with high-action scenes.
- Comparison: For tasks involving anatomy, physics, complex scenes, or accurate character generation, other models like Sora 2, Cling 2.5, and High Law O2 are preferred, often being cheaper and faster.
Where to Use VO 3.1
- Google Flow: Offers 100 free credits per month, allowing for approximately 5 videos using V3.1 Fast or 1 video using V3.1 Quality.
- Higsfield: A popular platform that supports VO 3.1 and other video models, with flexible image uploading.
- Chat LLM by Abacus AAI: Provides access to a wide range of AI models, including VO 3.1.
- Other Platforms: Hugging Face, Foul.ai, Waves, and Replicate also offer access to VO 3.1. Links to these platforms are provided in the video description.
The presenter concludes that while VO 3.1 is good for specific use cases like character consistency and audio generation, for most other needs, alternative models are superior due to their speed, cost-effectiveness, and better performance. The video encourages viewers to subscribe to the presenter's newsletter for more AI news.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development