Alibaba's new AI video tool WAN 2.5 Preview is WILD!

Prompt EngineeringAbout 3 min readSep 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Text-to-video model
  • Audio generation
  • Image-to-video
  • Speech synchronization
  • Video resolution (1080p)
  • AI video generation
  • Deep fakes
  • Multimodality
  • Deep alignment
  • Unified framework
  • Image animation
  • Driving video
  • Open source

One 2.5 Preview: Alibaba's New Text-to-Video Model

  • Introduction: The video introduces a new text-to-video model from Alibaba, One 2.5 Preview, which can natively generate audio. This is the second model from Alibaba with this capability, following V3.
  • Audio Synchronization: The model demonstrates impressive audio synchronization, especially with human characters.
    • Example: "Do you know one 2.5 preview? Yes, I come from the model."
  • Video Generation Capabilities:
    • Generates videos up to 10 seconds long.
    • Supports image-to-video generation.
    • Capable of generating videos at 1080p resolution.
  • Speech Synchronization Issues: Speech synchronization may have issues when animating characters.
  • Audio Effects: Audio effects are relevant to the scene.
    • Example: "Go on, get off my property. No visitors today or any day."
  • Complex Scene Generation: The model can generate complex scenes with accurate audio effects.
    • Example: "We use words like honor, code, loyalty. We use these words as the backbone of a life spent defending something. You use them as a punchline."
  • Camera Movements: The model supports complex camera movements.
    • Example: Music video scene with dynamic camera angles.
  • Video Stitching: Longer videos can be created by stitching together multiple 10-second clips, but this may result in glitches.
  • Input Methods: The model accepts a reference image with a text prompt or a direct text prompt to generate videos.
  • Sound Effects: The model generates relevant sound effects.
    • Example: UFO scene with audio effects that change based on the UFO's elevation.
    • Example: Video game scene where the character says "die, die" while firing, with added ammo effects.
  • World Understanding: The model demonstrates world understanding.
    • Example: Cat video with appropriate cat sounds at the right moment.
  • Availability: The model is available on file for testing.
  • Official Release: Alibaba officially released One 2.5 Preview, highlighting its new architecture and features.
    • Features: Native multimodality, deep alignment, and a unified framework for understanding and generation.
    • Supports input and output of text, image, video, and audio.

One Animate: Image Animation Based on Driving Video

  • Introduction: Alibaba released One Animate, which animates a specific image based on a driving video.
  • Animation Accuracy: The generated video's movements closely match the driving video.
    • Example: Animated character mimicking the movements and expressions from a driving video.
  • Open Source: One 2.2 Animate model is open source.
  • Custom Workflows: Custom workflows, such as Comfy UI, are available for animating images using an input video.

Implications and Concerns

  • Deep Fakes: The technology raises concerns about the ease of generating deep fakes.
  • Distinguishing Reality: It is becoming increasingly difficult to distinguish between real and fake videos.
  • New Possibilities: The technology opens up new possibilities that were previously unimaginable.

Call to Action

  • Sign up on One.video to test the new model.
  • A more detailed video will be created in the future.
  • Information on whether the model will be open source is forthcoming.

Conclusion

Alibaba's One 2.5 Preview and One Animate models represent significant advancements in AI-driven video generation. The ability to generate high-quality videos with synchronized audio and animate images based on driving videos opens up new creative possibilities. However, the technology also raises concerns about the potential for misuse, particularly in the creation of deep fakes. The presenter encourages viewers to explore the new models and consider the broader implications of generative media.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.