The most powerful AI video generator you can use OFFLINE

AI SearchAbout 6 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Vase by Alibaba, Wand 2.1, Comfy UI, Text-to-Video, Image-to-Video, Reference Video to Video, ControlNet, Pose Estimation, LoRA (Cosvid), UMT5XXL Text Encoder, VAE, Quantization, VRAM, GGUF, Workflow, Nodes, Bypassing Nodes, K Sampler, Steps, CFG, Seed, Inpainting, Outpainting, Reference to Video.

Installation and Setup of Vase

  • Vase Introduction: Vase is a free and open-source AI video generator by Alibaba, utilizing Wand 2.1. It's uncensored and allows for generating diverse content.
  • Full Version vs. Quantized Version: The full version of Vase requires 80GB of VRAM. A quantized version by Quanstack allows usage with as low as 8GB of VRAM. The presenter uses the Q6 version, requiring 14.5GB of VRAM, due to having a 16GB GPU.
  • Comfy UI Requirement: Comfy UI is necessary to run Vase. The presenter assumes the viewer has Comfy UI installed and refers to a previous video for installation instructions.
  • Downloading Vase Model: The Q6 version of the Vase model (vase_14B_Q6) is downloaded from the Quanstack Hugging Face page and placed in the ComfyUI/models/unet folder.
  • Workflow Download: An example workflow (V2V_example_workflow.json) is downloaded from the same Hugging Face page and saved in the ComfyUI folder.
  • Missing Nodes: Upon opening the workflow in Comfy UI, missing custom nodes are installed using the Comfy UI Manager. The Multi-GPU node may need to be manually installed.
  • Model Selection: The downloaded Vase model (Q6 version) is selected in the model selector node.
  • Low VRAM Configuration: If the GPU has low VRAM but the computer has high RAM, a specific node can be toggled to offset some computing to the RAM. The amount of VRAM on the GPU needs to be specified.
  • Cosvid LoRA: A Cosvid LoRA (Juan 2.1 Cosvid 14b Lura) is optionally downloaded to speed up generation. It's placed in ComfyUI/models/loras.
  • UMT5XXL Text Encoder: The correct UMT5XXL text encoder (Q6 encoder) is downloaded from the provided Hugging Face link and placed in ComfyUI/models/text_encoders.
  • VAE Download: The VAE (one 2.1) is downloaded from the provided link and placed in ComfyUI/models/vae.
  • Model Refresh: Pressing "R" in Comfy UI refreshes the dropdown lists to show the downloaded models.

Text-to-Video Generation

  • Bypassing Unused Nodes: For text-to-video, the reference image and reference video nodes are bypassed by selecting them, and pressing Ctrl+B.
  • Prompt Input: A text prompt is entered (e.g., "the girl is dancing in a sea of flowers slowly moving her hands").
  • Resolution Adjustment: The width and height of the output video are set (e.g., 1280x720). The noodles need to be disconnected to adjust the width and height.
  • Video Length: The length of the video is determined by the number of frames (e.g., 49 frames for approximately 3 seconds at 16 frames per second).
  • K Sampler Settings:
    • Steps: The number of iterations for video generation (20 is a good starting point, can be decreased to 15 for faster generation).
    • CFG: How closely the AI should follow the prompt (higher values are more literal, lower values allow for more creativity).
    • Sampler Name: The algorithm used for video generation.
    • Seed: The starting point for generation (fixed seed produces the same video, randomize for different videos).
  • Cosvid LoRA Settings: When using the Cosvid LoRA, the number of steps should be set to 4-6, and the CFG should be set to 1.
  • Output: The generated video is automatically saved in the ComfyUI/output folder.

Image-to-Video Generation

  • Unbypassing Image Node: The reference image node is unbypassed by selecting it and pressing Ctrl+B.
  • Image Upload: A reference image is uploaded.
  • Resolution Adjustment: The width and height of the output video can be adjusted. If the output resolution is smaller than the input image, the video will be cropped.
  • Prompt Input: A text prompt is entered (e.g., "she is talking").

Reference Video to Video Generation

  • Unbypassing Video Nodes: The reference video nodes are unbypassed.
  • Video Upload: A reference video is uploaded.
  • Resolution Adjustment: The width and height of the output video are set and connected to the K Sampler.
  • ControlNet Preprocessing: The input video needs to be pre-processed into a ControlNet video (e.g., pose map, depth map, edge map).
  • ControlNet Auxiliary Installation: Comfy UI ControlNet Auxiliary needs to be installed.
  • Pose Estimation: The default Canny pre-processor (edge detection) can be replaced with a pose estimator (e.g., OpenPose) to transfer movements based on poses.
  • Node Connection: The OpenPose node is connected to the image connector and the Control Video node.
  • Prompt Input: A text prompt is entered (e.g., "three cats are dancing").

Combining Reference Image and Reference Video

  • Unbypassing Image Node: The reference image node is unbypassed.
  • Image Upload: A reference image of the desired character is uploaded.
  • Video Upload: A reference video with the desired movements is uploaded.
  • Prompt Input: A text prompt is entered describing the desired scene (e.g., "the girls are dancing").

Examples and Demos

  • Dancing Cats/Bears: A reference video of ladies dancing is used to generate videos of cats or bears dancing with the same movements.
  • Bikini Women on the Beach: The same dance reference video is used to generate a video of women in bikinis dancing on the beach.
  • Surfing to Snowboarding: A video of a surfer is used to generate a video of a Japanese girl in a kimono snowboarding with the same poses.
  • Anime Girls Dancing: A reference image of anime girls is combined with a dance reference video to generate a video of the anime girls dancing.
  • Fight Scene: A reference image of a fight scene is combined with a boxing match video to generate a video of the characters in the image fighting with the same movements.
  • Cyberpunk City: A drone video of a city is transformed into a cyberpunk city at night with neon signs or an abandoned city overgrown with plants.

Limitations and Troubleshooting

  • Hand and Finger Imperfections: Imperfections in hands and fingers may occur, especially with low VRAM.
  • Quality Improvement: Increasing the number of steps, disabling the Cosvid option, or using a less quantized model may improve quality.
  • Error Troubleshooting: Users are encouraged to copy and paste error messages in the comments for troubleshooting assistance.

Nota AI Notetaker

  • Sponsor: Nota AI Notetaker is mentioned as a sponsor.
  • Features: Nota automatically transcribes, summarizes, and organizes spoken content.
  • Integration: It supports platforms like Zoom, Google Meet, Teams, and Webex.
  • Use Cases: Useful for consultants, sales reps, customer support, and students.

Conclusion

Vase is a powerful and flexible AI video generator that allows for text-to-video, image-to-video, and reference video to video generation. The combination of reference images and videos provides a high degree of control over the output. The quantized versions make it accessible to users with lower VRAM GPUs. While there are some limitations, the tool offers significant creative potential.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.