Veo 3.1 is HERE - Full Test vs Sora 2 Pro & Wan 2.5 (Higgsfield AI)
By Futurepedia
Key Concepts
- AI Video Generation Models: VO 3.1, Sora 2 Pro, Juan 2.5, Cling 2.5, Hyo 2, Seedance Pro.
- Higsfield: A platform for testing and comparing AI models.
- Image-to-Video: Generating video from a static image.
- Text-to-Video: Generating video from a text prompt.
- Physics Simulation: Testing AI's understanding of physical interactions.
- Audio Generation: AI's capability in generating sound effects, dialogue, and music.
- Lip Sync: The synchronization of mouth movements with spoken dialogue.
- Character Consistency: Maintaining a character's appearance across multiple shots.
- Motion Graphics: Animated visual elements used in media.
- Start/End Frames (VO 3.1): Feature allowing users to define the beginning and end of a generated video.
- Ingredients Feature (VO 3.1): Feature allowing users to combine multiple images into a video.
- Camera Control Presets (Juan 2.5): Predefined camera movements for video generation.
- Sketch-to-Video (Sora 2): Generating video from a simple sketch.
- Prompt Engineering: Crafting effective text prompts for AI models.
AI Video Model Comparison: VO 3.1, Sora 2 Pro, and Juan 2.5
This analysis compares the capabilities of three prominent AI video generation models: VO 3.1, Sora 2 Pro, and Juan 2.5, alongside other models like Cling 2.5, Hyo 2, and Seedance Pro, across various use cases. The testing was conducted using the Higsfield platform, which allows for simultaneous generation and comparison across models, saving time and avoiding multiple subscriptions.
Physics and Motion Tests
1. Domino Falling Challenge:
- Objective: To generate a video of a person tipping over dominoes, causing them to fall.
- Results:
- VO 3.1, Juan 2.5, Cling 2.5, Hyo 2, and Seedance Pro all failed to accurately simulate the physics, producing "weird stuff" or being "total fails."
- Sora 2 Pro was the only model that successfully generated this scenario, which was a significant achievement as it had proven difficult for all previous models.
2. Basketball Shot:
- Objective: To generate a video of a basketball shot going into the hoop using text-to-video.
- Results:
- Sora 2 Pro: Missed the shot despite using "sort of accurate physics," resulting in a "weird looking generation."
- VO 3.1: The ball went in the hoop. The model tends to use slow motion, which can be improved by speeding up playback (150-200%). Sound effects synced perfectly. Considered a "big step up" from previous models.
- Juan 2.5: The ball went in the net but got "caught too long at the bottom of the net." Also felt slow-motion, but speeding it up made it "pretty solid."
- Conclusion: None were perfect, but all showed significant improvement over older models, with the potential for success through rerolls.
Audio and Dialogue Capabilities
1. Sound Effects and World Understanding (Alarm Clock Test):
- Objective: To test sound effects, basic text, and world understanding with an alarm clock scenario.
- Results:
- Juan 2.5: "Almost perfect." Correctly depicted time switching and the alarm going off, but added an extra "AM." The character's expression was right, but he covered his head instead of hitting snooze.
- VO 3.1: Numbers were less accurate, and sound effects were not as good. "Definitely not bad."
- Sora 2 Pro: "Basically perfect," with the alarm starting slightly too soon being the only minor flaw. It also offered multiple shots and maintained visual consistency. The annoyance and slamming action were well-captured with perfect sound sync.
2. Dialogue Generation (Text-to-Video):
- Scenario: A woman introducing her shark friend, "Fluffy."
- Sora 2 Pro: "Looks pretty real" but the mouth moved "under the water as she's speaking."
- VO 3.1: High detail in the scene, potentially "too high quality" for a handheld shot. Lip sync failed halfway through.
- Juan 2.5: "Basically perfect," "knocked that one out of the park."
3. Dialogue Generation (Image-to-Video):
- Scenario: A Viking persona speaking.
- Sora 2 Pro: Refused to generate due to realistic human faces, even AI-generated ones.
- Juan 2.5: "Pretty weird on the facial expressions" and "lip sync was just completely off."
- VO 3.1: "Just perfect." Added motion, maintained consistent facial features, and coherent overall motion for multiple characters. This was the "absolute winner" for this category.
4. Action and Dialogue:
- Scenario: A character stating, "Next time, I'm taking the front door."
- Juan 2.5: Physics issues, including going through a rail and a "sort of okay landing." Lip sync was absent at the beginning but synced at the end. "Not particularly good."
- VO 3.1: "Super weird on the physics." A character disappeared and reappeared. Dialogue at the end was "solid."
- Sora 2 Pro: "Really, really good." Running looked great, with quick cuts like an action scene. Added two jumps, and missed a small piece of dialogue. Considered the "winner out of these."
5. Singing:
- Scenario: Singing lyrics about lost love.
- Juan 2.5: Generated the video and lyrics correctly, but the singing was "definitely not good."
- Sora 2 Pro: Banjo playing looked "off," and singing was "not great."
- VO 3.1: "Far above the other options." A "really good sounding song" with a great voice. Strumming was passable. This was a surprising high-quality generation.
6. Full Band Performance:
- Scenario: A high-energy pop punk band performing about heartbreak from being dumped by an AI.
- Juan 2.5: "Terribly." Poor visuals, "terrible" song, and nonsensical lyrics. Juan tends to generate random noises if not specified.
- Sora 2 Pro: Lyrics made sense, song was "okay." Had the "vibe" of a live performance, but motion felt "pretty off."
- VO 3.1: "Amazing." Higher fidelity, solid camera work, and perfect mouth sync. Jump and headbanging motions synced with the song.
7. Deep Conversation:
- Scenario: Two people having a deep conversation about life.
- Juan 2.5: Generated a "bad commercial about life" with random noises, as dialogue was not specified.
- VO 3.1: "Really good." Natural speech, turn-taking, and multiple characters. Minor "weirdness here and there," but "pretty solid."
- Sora 2 Pro: "Nailed it, too." Felt good, but one character spoke "a little too fast." Good back-and-forth dialogue, including brief overlapping speech.
8. Argument Scene:
- Scenario: An argument between characters (dialogue not specified).
- Juan 2.5: Looked and felt good, but the words were "just nonsense."
- VO 3.1: Dialogue was good, but felt "super off," like a "soap opera satire." The wrong character said the wrong line at one point.
- Sora 2 Pro: "Just incredible." Nailed emotions, felt authentic, and had perfect lip-sync. Sora is improving in conveying emotions.
9. Comedian Telling a Joke:
- Scenario: A comedian telling a joke on stage.
- Sora 2 Pro: Spoke "super fast," and the joke was not clear. Visuals were good, but the delivery was poor.
- VO 3.1: Visuals looked good, but the joke "didn't even make any sense."
- Conclusion: Neither model was particularly funny.
10. Detective Conversation:
- Scenario: An intense conversation between two detectives, with the captain involved.
- Juan 2.5: "Really good." The vibe and aesthetics of the diner and shot angle were praised. It was intended as a single speaker but worked well with two.
- VO 3.1: Added unwanted music, which hinders merging multiple scenes. Aesthetic was not preferred. Lip-sync was solid and delivery was good.
- Sora 2 Pro: Shot was "really, really good." Added music (though not requested) which was better than VO's choice. Dialogue was good, and the intensity and eye contact felt appropriate.
11. Potato Podcast:
- Scenario: A potato speaking on a podcast about potatoes.
- Sora 2 Pro: "Super weird looking potato." Lip-sync worked well.
- VO 3.1: "Way better." Felt more like a podcast, with a better studio setting and potato appearance. "Basically, just perfect generation."
- Juan 2.5: Used a "super weird style choice." Lip-sync was "off throughout the shot." Even with a Midjourney image, lip-sync remained an issue for non-human characters.
Unique Features and Capabilities
1. VO 3.1 - Start and End Frames:
- Description: A new feature allowing users to define specific start and end frames for generated videos.
- Application: Extremely useful for AI filmmaking, enabling shots that were previously difficult or impossible.
- Example: Generating a dialogue scene with a specific man and woman, including a "J cut" where the second speaker starts before the visual cut. This feature was essential for creating shots in a short film for the Skill Leap course.
2. VO 3.1 - Ingredients Feature:
- Description: Allows users to combine up to three images into a single video.
- Application: Useful for creating consistent characters across scenes without needing separate image generators.
- Example: Combining a grid of eight items with an image of a marble room to create a video depicting specific characters wearing certain items in that setting. The model understood all details, including the woman in a hat riding a tiger and a man in a colored jacket riding a capybara wearing sunglasses.
3. Juan 2.5 - Camera Control Presets:
- Description: Offers various camera control presets, including "handheld car grip" and "bullet time."
- Application: Helpful for generating complex camera movements that are difficult to prompt.
- Example: Using the "eyes in" preset with an image of a witch and tigers, resulting in a detailed shot focused on her eye.
4. Sora 2 - Sketch-to-Video:
- Description: Converts low-quality sketches into multi-shot videos.
- Application: Popular for creating advertisements.
- Example: Creating a video for the "Skill Leap" platform by drawing a sequence of sketches depicting a person at a computer, the screen, a relaxed pose, and a final screenshot with text. Despite the "terrible sketch," the AI generated a coherent video with much of the text included.
Style and Specificity Tests
1. Group of Friends Laughing:
- Objective: To generate a scene where friends laugh after someone says something funny.
- Results:
- Juan 2.5: "Pretty good job." Faces remained consistent, and the laughter appeared genuine.
- VO 3.1: Looked great but tended towards "stock footage looking shots."
- Winner: Sora (implied, as VO's specific example was about a vet and parrot, not the friends laughing).
2. Woman Crying:
- Objective: To generate a woman crying.
- Results:
- Juan 2.5: "Almost perfect." Genuine emotion and realistic crying sounds.
- VO 3.1: Tears looked "terrible," like "wax dripping out of her eye," making the scene unsettling.
- Sora 2 Pro: "Really, really good" at conveying emotion. However, it added unwanted dialogue and music, which can "ruin" shots.
3. Text Generation (Channel Intro):
- Objective: Creating a channel intro with text and circuit board graphics.
- Results:
- VO 3.1: Started well but had a "weird" backward motion. Added unwanted cuts and failed to complete the text and circuit board convergence. "Not a great job."
- Winner: VO 3.1 was "by far the best in every way" for this specific intro.
4. Cursive Text on Chalkboard:
- Objective: Writing "hello" in cursive on a chalkboard.
- Results:
- Juan 2.5: Drew with chalk decently but added "random lines."
- Sora 2 Pro: Followed the chalk path perfectly but did not write "hello," only flashing it at the end.
- VO 3.1: Drew in the wrong spot and didn't follow the chalk path well, but had the best sound effects and ended with the word "hello."
- Conclusion: None of the models performed this task "too well."
5. Restaurant Pan Shot:
- Objective: A pan through a restaurant showing conversations.
- Results:
- Juan 2.5: Faces remained coherent. A character in the back disappeared. "Good job following the camera movement."
- Sora 2 Pro: Tried to make it an ad, with "weirdness in the faces." A character talked to an empty seat. "Not bad. Not great."
- VO 3.1: Started well but focused on one table, then characters disappeared "out of nowhere." "Super weird and random total fail."
6. Inside a Fridge Shot:
- Objective: A shot from inside a fridge.
- Results:
- Juan 2.5: Physics of bottles looked "a bit off." Understood the shot type.
- Sora 2 Pro: "All sorts of crazy physics," couldn't get it right.
- VO 3.1: Didn't understand the shot but produced an "all right shot for what it did end up doing."
- Conclusion: "Fail all around on that one."
7. Pixar Style Animation (Dragon and Fire):
- Objective: A dragon blowing fire to heat a cup, followed by sipping and smiling.
- Results:
- Juan 2.5: "Basically perfect." Good sipping action and sound effects. "Great job all around."
- Sora 2 Pro: "Weird puff of fire" and "strange sip." Face-smashing action. "Not very good."
- VO 3.1: "Weird sound effects" at the beginning, but otherwise a good shot with multiple actions.
- Winner: Juan 2.5.
8. Unique Midjourney Style:
- Objective: Maintaining a specific Midjourney style with a character scooping salmon.
- Results:
- Juan 2.5: "Pretty good job." Salmon style was not perfect but consistent. Splashes were solid.
- Sora 2 Pro: Didn't fully follow the prompt, but style remained consistent, though definition was lost.
- VO 3.1: "Weird choice" with the salmon popping up. Kept style well, but the salmon disappeared at the end.
- Winner: Juan 2.5.
9. Tiger Eating Ramen and Revealing Panda:
- Objective: A tiger eating ramen, then panning to a panda eating at the same table.
- Results:
- Juan 2.5: "Pretty good job" on eating physics. Panda fit well in the scene. However, dialogue was weird and unsolicited, with the tiger speaking while eating.
- Sora 2 Pro: Panda looked "all right," but the tiger's dialogue was "super off" and unwanted.
- VO 3.1: "Basically perfect." Good bite from the tiger, smooth pan to the panda, and the panda took a bite. Panda fit well in the scene.
Difficult Physics and Motion Challenges
1. Threading a Needle:
- Objective: Getting a tiny piece of thread through the eye of a needle.
- Results: No model has been able to achieve this successfully. They often appear close but fail at the last moment or become "weird."
2. Salsa Dancing:
- Objective: Generating salsa dancing.
- Results: All models captured the "vibe" of salsa dancing, but close inspection revealed physics inaccuracies and "morphing and weird stuff."
3. Break Dancing:
- Objective: Generating break dancing with a large crowd.
- Results: Models captured the "vibe" and movements, feeling coherent even with physics errors. However, no model has done a "great job" yet.
4. Man Riding a Horse on Another Horse:
- Objective: A classic prompt testing prompt understanding.
- Results:
- Juan 2.5: Made a "very strange decision" but arguably got it "right" in terms of a realistic outcome.
- Sora 2 Pro: Generated what was likely the intended visual.
- VO 3.1: Made a "good attempt" but used perspective to create the illusion.
- Winner: Juan 2.5 for accuracy to reality.
5. Glass Breaking with Fluid Dynamics:
- Objective: Simulating fluid dynamics and glass breaking.
- Results:
- Juan 2.5: "Definitely not a good job."
- Sora 2 Pro: Water looked okay, but the glass didn't break.
- VO 3.1: Shattered the glass and added ice, with a "crazy amount" of water.
- Conclusion: "None of them got even close on that one."
6. Samurai Fight Scene:
- Objective: A samurai fight scene.
- Results:
- Juan 2.5: Looked like "weird stock footage in a studio." Swords moved roughly correctly.
- Sora 2 Pro: Looked like a "video game" with motion and camera work. Swords disappeared and reappeared with warping. Camera movement was praised.
- VO 3.1: "Feels awful. Super strange."
- Winner: Sora 2 Pro, despite warping, due to camera movement.
7. Clown Juggling on a Unicycle:
- Objective: A clown juggling while riding a unicycle through a carnival.
- Results:
- Juan 2.5: Unicycle looked okay, but juggling failed with balls appearing and disappearing.
- Sora 2 Pro: Pins went "all over the place."
- VO 3.1: A ball disappeared early. Felt most like accurate physics, but the character was not riding a unicycle and had handlebars.
- Conclusion: "Not even close," but VO 3.1 was the closest.
Overall Summary and Conclusion
Clear Winners by Category:
- Music: VO 3.1 (by a significant margin).
- Unique Styles: Juan 2.5 (based on limited tests).
- Dialogue: Sora 2 Pro won some, Juan 2.5 won one, and VO 3.1 won a couple. VO 3.1 was the winner for non-human dialogue.
- Emotions: Sora 2 Pro won, but it was a close call.
- Text: No clear winner; depended on the situation. VO 3.1 was best for title screens.
- Physics Tests (e.g., dominoes, glass breaking): Sora 2 Pro, with close competition.
- General Complex Movements (dancing, fighting, running): Sora 2 Pro won across the board.
Overall Winner: Sora 2 Pro was the overall winner by far based on the tally of categories.
Massive Caveats and Considerations:
- Sora 2 Pro's Censorship: Sora 2 Pro cannot generate realistic human faces when using image-to-video, severely limiting its use for many scenarios.
- VO 3.1's Advanced Features: VO 3.1's start/end frame and ingredients features are crucial for specific filmmaking needs and enable shots impossible with other models.
- Juan 2.5's Versatility and Less Censorship: Juan 2.5 is generally less censored and works well for many dialogue scenes.
- Other Models: Cling 2.5 and Hyo 2 are also noted for complex motion and camera movements.
- Higsfield Platform: The Higsfield platform is highly recommended for its ability to test and compare multiple AI models in one place, saving significant time and effort.
Final Takeaway: While Sora 2 Pro demonstrated strong performance in many areas, its limitations with realistic human faces and censorship make it unsuitable for all tasks. VO 3.1's unique features and strong performance in music and specific dialogue scenarios make it a valuable tool. Juan 2.5 offers good versatility and less restriction. The choice of model ultimately depends on the specific use case and desired outcome.
The video concludes by highlighting the importance of using the right tool for the job and mentions Futuredia's comprehensive AI course platform for those wanting to delve deeper into AI.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Don't use Fable 5 in Claude… do this instead
David Ondrej

Claude Code VS Codex VS GLM VS Kimi: I Tested EVERY AI Coding Agent.. Here's my thoughts.
AICodeKing

OpenAI vs. Anthropic: The AI Vibe Shift Explained #shorts
Authority Hacker Podcast

The Best Model For AI Coding Is...
corbin

Is Kling 3.0 Actually the Best? Full Breakdown vs Competition
Futurepedia

Was I Wrong? GPT 5.3 vs Opus 4.6 — Round 2
Eduards Ruzga

NEW GPT Image 1.5 vs Nano Banana Pro
AI Search