How to Make Consistent Characters in Veo 3 (AI Video Tutorial)

AI Video SchoolAbout 4 min readJun 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Consistent character generation in Runway V3
  • Prompt engineering using Whisk and Gemini
  • Voice cloning with 11 Labs
  • Caption removal using Runway inpainting and CapCut AI remove
  • Runway V3 fast vs. quality modes

Consistent Character Generation in Runway V3

The speaker details their process for creating consistent characters in Runway V3, acknowledging the current lack of character reference image support. The foundation lies in leveraging detailed prompt engineering.

Prompt Engineering with Whisk and Gemini

  1. Image Generation in Whisk: The speaker initially created an image of the character in Google's Whisk using a specific prompt. This image served as the basis for extracting a more robust prompt.
  2. Subject Analysis in Whisk: The image was then analyzed within Whisk by dragging it to the "subject" area. This generated a detailed image description, providing insight into how the AI perceives the character.
  3. Prompt Refinement with Gemini: The original Whisk prompt and the image description generated by Whisk were fed into Gemini (Google's AI model). The speaker instructed Gemini to create a detailed V3 description of the man, focusing on his face and ignoring wardrobe, to be used as a consistent template.
  4. Character Naming and Voice Prompting: Gemini was also used to suggest a name for the character (Aram) and to generate voice prompt options for V3, aiming for a consistent voice profile.
  5. Core Prompt Development: The speaker requested core prompts for Aram's visual description, voice, and a cinematic 35mm style. This resulted in a set of prompts that could be consistently used across different generations.
    • Example: The speaker provides an example of a full prompt including the physical description of Aram and how he speaks.
    • Example: The speaker also provides an example of a prompt in a high-end documentary style where Aram says, "I used to wake up and worry about all of the things I needed to do that day."
  6. Key Takeaway: The core strategy is to build a very detailed description of the character and consistently use that description in all prompts.

Caption Removal Techniques

The speaker addresses the issue of unwanted captions generated by Runway V3.

  1. Gemini Experiment: The speaker initially attempted to avoid captions by generating videos directly within Gemini, based on a Discord suggestion. However, this method proved inconsistent, sometimes producing captions.
  2. Runway Inpainting: The speaker then explored AI-powered caption removal methods, starting with Runway's inpainting tool. This involved using a brush to paint over the captions, effectively erasing them. The speaker notes that while this method works, subtle artifacts might be visible upon close inspection.
  3. CapCut AI Remove: The speaker found CapCut's AI remove feature to be more effective. By selecting the video on the timeline, navigating to the "AI remove" option under the "video" tab, and using the quick brush tool, the captions could be removed with smoother results. The speaker highlights the impressive performance of CapCut's AI in seamlessly removing captions even in areas with intricate details like a collar with lines.

Voice Cloning and Consistency with 11 Labs

The speaker encountered issues with Aram's voice, particularly unwanted accents, possibly due to the character's description as an Armenian farmer.

  1. Voice Cloning with 11 Labs: To address this, the speaker used 11 Labs to clone Aram's voice. This involved extracting approximately 10 seconds of audio from existing clips of Aram speaking.
  2. Voice Changing Experiment: The speaker then attempted to use 11 Labs' voice changer on all the generated audio, hoping to eliminate the accent and achieve a consistent voice. However, this method was not entirely successful, as some videos still retained an accent.
  3. Text-to-Speech Substitution: As a workaround, the speaker resorted to using text-to-speech in 11 Labs for the problematic videos, manually adjusting the timing to match the original audio.
  4. Rationale: The speaker emphasizes the importance of both a consistent face and a consistent voice in selling the illusion of a real character.

Runway V3 Fast vs. Quality Modes

The speaker discusses the different generation modes in Runway V3 and their impact on results.

  1. Credit Cost: Initially, each V3 video cost 100 credits. Now, a "fast" mode is available at 20 credits per video.
  2. Unexpected Results: The speaker notes that the fast mode can sometimes produce surprisingly good results, even rivaling the quality mode. Additionally, the fast mode can generate unexpected and surreal visuals.
    • Example: The speaker mentions that one of their favorite shots in the entire documentary was generated using the fast mode.
  3. Mode Selection: The speaker demonstrates how to switch between the V3 fast (20 credits) and V3 quality (100 credits) modes within the Runway interface by clicking the settings button.

Conclusion

The speaker provides a detailed walkthrough of their process for creating consistent characters in Runway V3, highlighting the importance of detailed prompt engineering, AI-powered caption removal, and voice cloning techniques. They also share insights into the different generation modes available in Runway V3. The key takeaway is that achieving consistency requires a combination of careful planning, creative use of AI tools, and a willingness to experiment and adapt.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.