FINALLY! Open-source AI video with audio!

By AI Search

Share:

Key Concepts

  • OVI: An open-source AI video generator with built-in audio capabilities, allowing for video creation from text prompts.
  • Comfy UI: A popular, performance-optimized interface for running open-source AI models locally, known for its node-based workflow and VRAM management features.
  • VRAM (Video RAM): Dedicated memory on a GPU, crucial for running AI models; OVI has specific VRAM requirements.
  • CPU Offload: A technique used in Comfy UI to offload excess memory from the GPU's VRAM to the system's RAM, enabling models to run on systems with lower VRAM.
  • Text-to-Video: The process of generating video content based on a textual description or prompt.
  • Image-to-Video: The process of generating video content that starts from a provided reference image.
  • Dialogue Tags ([start]...[end]): Specific syntax used within OVI prompts to delineate spoken dialogue.
  • Audio Tags ([oddcap]...[oddcap]): Specific syntax used within OVI prompts to describe desired voice characteristics or background audio.
  • Loras (Low-Rank Adaptation): A method for fine-tuning AI models to generate specific styles, characters, or actions without extensive retraining, a benefit of OVI's open-source nature.
  • Quantized/Compressed Models (GGUFS): Smaller, more memory-efficient versions of AI models, anticipated for OVI to enable broader accessibility on lower-spec hardware.
  • Seed: An initial value that determines the starting point of random noise in AI generation, used to ensure reproducibility or generate variations.
  • Sampling Steps: The number of iterations an AI model performs during generation; more steps generally lead to higher quality but slower output.
  • Negative Prompts: Textual descriptions of elements or qualities that users wish to exclude from the AI-generated video or audio.

OVI: The First Open-Source AI Video Generator with Built-in Audio

The video introduces OVI, a new open-source AI video generator that uniquely integrates audio capabilities, including dialogue, background audio, and sound effects, directly from text prompts. Positioned as a free, offline, and unlimited alternative to models like V3 or Sora 2, OVI offers significant flexibility and control over video creation.

Core Features and Capabilities

OVI's robust feature set allows for detailed and expressive video generation:

  • Prompt-Driven Generation: Users define video content, dialogue, voice characteristics, and background audio through structured text prompts.
    • Dialogue Specification: Dialogue is precisely defined by enclosing it within [start] and [end] tags (e.g., [start]Come here. Come here.[end]).
    • Voice and Background Audio Control: Specific voice qualities or background sounds are indicated using [oddcap] tags.
  • Advanced Animation and Realism: OVI excels in generating realistic lip-syncing and animating the entire body, including nuanced hand gestures, to align with the context and dialogue.
  • Complex Scene Handling: The model supports multiple characters, allows for specifying different actions within a scene, and can generate sequences of actions and dialogues.
  • Multilingual Support: OVI demonstrates the ability to generate dialogue in various languages, with examples shown in English, German, and Korean.
  • Contextual Sound Effects: Beyond spoken words, OVI can generate appropriate sound effects based on the video's content, such as rain noise.
  • Emotional and Expressive Range: It effectively handles diverse expressions and emotions in generated characters.
  • Image-to-Video Functionality: A standout feature is the capability to use a reference image as the starting frame for a video. Crucially, this works even with images containing realistic people, a functionality not available in Sora 2 for such content.
  • Singing Generation: OVI also possesses the unique ability to generate singing performances.

Comparative Analysis and Advantages

While the current quality of OVI, particularly for complex movements like gymnastics, is noted as "not as good as like VO3 or Sora 2," its open-source nature provides substantial long-term advantages:

  • Extensibility via Loras: Being open-source, OVI can be enhanced with Loras (Low-Rank Adaptation) to fine-tune for specific characters, styles, or actions, including potentially uncensored content, which is a limitation for closed-source alternatives like V3, Sora 2, and 1 2.5.
  • Community-Driven Improvement: The open-source community is expected to rapidly develop more optimized and higher-quality versions, including quantized or compressed models (GGUFS), making OVI accessible on even lower VRAM systems in the near future.

Online Access Options

For users without the necessary local hardware, OVI can be accessed via several online platforms:

  • waves.ai: Offers both image-to-video and text-to-video generation. Users receive $1 in free credits, equating to approximately six free generations (at 15 cents per generation).
  • foul.ai: A paid service providing text-to-video and image-to-video options, costing 20 cents per generation.
  • replicate: Currently supports only image-to-video generation, priced at roughly 20 cents per generation.
  • Hugging Face: Provides an accessible interface for interacting with OVI.

Comprehensive Local Installation and Usage Guide with Comfy UI

The video provides a detailed, step-by-step tutorial for installing and running OVI locally using Comfy UI, which is highly recommended for its user-friendliness, performance optimizations, and efficient memory management through CPU offload.

System Requirements:

  • GPU: A CUDA-compatible GPU is essential.
  • VRAM: While 24 GB of VRAM is officially recommended, successful operation has been reported on systems with as low as 16 GB of VRAM.
  • Memory Management: Even with 24 GB VRAM, enabling CPU offload is crucial to prevent out-of-memory errors by utilizing system RAM for excess memory.

Installation Steps:

  1. Comfy UI Prerequisite: Ensure Comfy UI is already installed (the Windows portable version is strongly advised).
  2. Navigate to Custom Nodes: Open the Comfy UI/custom_nodes directory.
  3. Clone Repository: In the custom_nodes folder, open a command prompt (cmd) and execute git clone [ComfyYOV GitHub URL]. The correct URL should be copied from the GitHub repository's green "Code" button.
  4. Change Directory: Navigate into the newly cloned ComfyYOV folder using the command cd ComfyYOV.
  5. Install Dependencies: Run pip install -r requirements.txt to install all required Python packages. This process may take a considerable amount of time due to the extensive list of dependencies.

Additional Model File Downloads: Two specific model files must be downloaded manually before running OVI:

  1. Text Encoder: Download the FPH version (approximately 7 GB) for systems with lower VRAM. Save this file to Comfy UI/models/text_encoders.
  2. VAE File: Download the 1 2.2 VA file (approximately 1.4 GB). Save this file to Comfy UI/models/VAE.
    • Note: The primary OVI models (FP8/BF16) and the audio generator model will be automatically downloaded during the first execution of the OVI workflow.

Comfy UI Workflow Setup and Configuration:

  1. Launch Comfy UI: Start the Comfy UI application.
  2. Load Workflow: Drag and drop the workflow.json file, located in Comfy UI/custom_nodes/ComfyYOV/workflow_example, onto the Comfy UI interface.
  3. Configure Workflow Nodes:
    • OV Model Selection: Choose between the FP8 version (suitable for 16-24 GB VRAM) or the full BF16 model (for systems with over 32 GB VRAM).
    • CPU Offload Model: Set this parameter to true to enable memory offloading, which is vital for preventing out-of-memory errors on systems with limited VRAM.
    • Device: Specify the GPU device number, typically 0 for a single GPU setup.
    • VAE Load: Select the 1 2.2 VA file from the dropdown.
    • Text Encoder Load: Select the appropriate text encoder (FP8 for lower VRAM, BF16 for higher VRAM).
    • Accelerator Methods: Generally, leave this set to auto; options include flash attention or Sage Attention.
    • Prompt Input: Enter the detailed video description, including dialogue (using [start]...[end] tags) and voice/audio specifications (using [oddcap]...[oddcap] tags). Referencing the original OVI GitHub repository for prompt examples is recommended.
    • Video Dimensions: Define the height and width of the final video output.
    • Seed: Set to randomize for unique generations each time, or a fixed number for reproducible results.
    • Solver Name: Typically left at its default setting.
    • Sampling Steps: Adjust this value to balance generation quality and speed; more steps generally yield higher quality but slower generation.
    • Negative Prompts: Input any elements or qualities that should be excluded from the generated video or audio.
    • Image-to-Video Integration: If using an image as a reference, connect the image upload node accordingly.
    • Frame Rate: The default frame rate is 24 frames per second.
    • Save Output: Toggle this setting to true to ensure the generated video is saved to the designated output folder.

Demonstrations: The video includes practical demonstrations of both text-to-video and image-to-video generation. The text-to-video example illustrates the automatic download of the necessary audio and FP8 models during its first run, while the image-to-video example showcases generation from an uploaded image with a custom prompt and specified voice.

Sponsor Spotlight: LTX Studio

The video features LTX Studio, an all-in-one platform designed to streamline the entire video workflow, from initial storyboarding and shot planning to creating professional final videos. Key offerings include:

  • Their proprietary open-source LTXV model.
  • Integrations with Nano Banana for photo editing and Flux Premium for generating high-quality images.
  • A new text-to-speech feature powered by Google Gemini 2.5 Pro, supporting multiple languages, accents, and emotional control.
  • Training on licensed Getty Images and Shutterstock datasets, which allows for free commercial use for most businesses.

Synthesis and Conclusion

OVI marks a significant milestone as the first open-source AI video generator with integrated audio. While its current generation quality may not yet rival leading closed-source models, its open-source nature fosters rapid community development. This allows for future enhancements like Loras for custom content and the development of more optimized, lower-VRAM versions. The presenter anticipates substantial improvements in open-source video models with audio in the coming months and encourages users to experiment with OVI and seek support for any installation issues.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video