But how do AI videos actually work? | Guest video by @WelchLabsVideo

3Blue1BrownAbout 5 min readJul 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Diffusion models, Brownian motion, CLIP (Contrastive Language-Image Pre-training), latent space, embedding space, denoising, DDPM (Denoising Diffusion Probabilistic Models), score function, time-varying vector field, stochastic differential equation, Fokker-Planck equation, DDIM, flow matching, classifier-free guidance, negative prompts, conditioning, cross-attention.

1. Introduction: AI, Diffusion, and Physics

AI systems excel at generating videos from text prompts, a process deeply connected to physics. Diffusion models, central to this generation, are remarkably equivalent to Brownian motion, but with time reversed and in high-dimensional space. This connection provides algorithms for image and video generation and offers intuitions for model operation.

Example: The video showcases an astronaut generated by the open-source model WAN 2.1, demonstrating the ability to add elements like a flag or laptop via text prompts.

2. Hands-on with a Diffusion Model (WAN 2.1)

The video generation process in WAN 2.1 starts with random noise.

Process:

  1. Initialization: A random number generator creates a video of pure noise.
  2. Transformation: The noise video is fed into a transformer (similar to those used in large language models like ChatGPT).
  3. Iteration: The transformer outputs a slightly structured video, which is added back to the original noise video. This process repeats, refining the video over 50 iterations.

3. Three Parts of Diffusion Models

The video breaks down diffusion models into three key areas:

  1. CLIP (Contrastive Language-Image Pre-training): A model that learns a shared space between words and pictures.
  2. Diffusion Process: The process of adding and removing noise, with a focus on the connection to physics.
  3. Combining CLIP and Diffusion: How CLIP is used to guide the diffusion process based on text prompts.

4. CLIP: A Shared Embedding Space for Images and Text

CLIP, released by OpenAI in 2021, consists of two models: a language model and a vision model. Both models output vectors of length 512. The goal is to make the vectors for matching image-caption pairs similar.

Training Approach:

  1. A batch of image-caption pairs is passed through the image and text models.
  2. The similarity between matching pairs is maximized, while the similarity between non-matching pairs is minimized.
  3. Cosine similarity is used to measure the similarity between vectors.

Example: Taking the difference between the CLIP vector of "me wearing a hat" and "me not wearing a hat" results in a new vector that closely corresponds to the word "hat."

Key Point: CLIP creates a "vector space of pure ideas," allowing mathematical operations on concepts within images and text.

5. Diffusion Models: From Noise to Images

Diffusion models generate images by transforming pure noise into realistic images step by step. The core idea is to add noise to training images until they are completely destroyed, then train a neural network to reverse this process.

Naive Approach (Doesn't Work Well): Training the model to remove noise one step at a time (predicting step n-1 from step n).

Berkeley Team's Approach (DDPM):

  1. Noise Addition During Generation: Random noise is added at each step of the image generation process.
  2. Predicting Total Noise: The model is trained to predict the total noise added to the original image, skipping intermediate steps.

Example: Removing the noise addition step in Stable Diffusion 2 results in a blurry image.

6. Diffusion as a Time-Varying Vector Field

Diffusion models can be understood as learning a time-varying vector field.

Simplified Example (2D Toy Data):

  • Images are represented as points in a 2D space (two pixels).
  • Adding noise is equivalent to taking a random step.
  • The model learns to reverse these random walks, pointing back to the original data distribution (e.g., a spiral shape).

Score Function: The learned direction pointing back towards the original data distribution.

Time Conditioning: The model is conditioned on a time variable (t) representing the number of steps taken in the random walk. This allows the model to learn coarse vector fields for large t and refined structures as t approaches 0.

7. Resolving the Mystery of Noise Addition

Adding random noise during image generation leads to sharper images because it prevents all points from collapsing to the mean or average of the dataset.

Mathematical Explanation: The model learns to point to the mean of the dataset, conditioned on the input point and time. Adding noise allows the model to sample from a Gaussian distribution around this mean.

8. DDIM: Faster Image Generation

DDIM (Denoising Diffusion Implicit Models) allows for high-quality image generation without adding random noise during the generation process, significantly reducing the number of steps required.

Key Idea: Using an ordinary differential equation (ODE) instead of a stochastic differential equation (SDE) to govern the diffusion process.

Result: DDIM generates high-quality images deterministically in fewer steps, without requiring changes to model training.

9. Guiding Diffusion with CLIP: Text Prompts

CLIP's shared representation of images and text can be used to guide the diffusion process.

Process:

  1. A text prompt is passed into the CLIP text encoder to generate an embedding vector.
  2. This embedding vector is used to steer the diffusion process towards the image or video described by the prompt.

Conditioning: Passing the text vector as another input into the diffusion model during training.

10. Classifier-Free Guidance: Enhancing Prompt Adherence

Classifier-free guidance leverages the differences between an unconditional model (not trained on a specific class) and a model conditioned on specific classes.

Process:

  1. Train a diffusion model with and without class information.
  2. Subtract the output of the unconditioned model from the output of the conditioned model.
  3. Amplify the resulting vector and use it to guide the diffusion process.

Negative Prompts: Specifying features to avoid in the generated video and subtracting the resulting vector from the model's output.

11. Conclusion

Diffusion models have progressed rapidly, leading to impressive text-to-video models. The ability to combine trained text encoders with complex diffusion processes is remarkable. These models represent a new class of machine, requiring only language to create lifelike images and videos. The guest video was created by Stephen Welch of Welch Labs.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.