Stanford CS231N Deep Learning for Computer Vision| Spring 2025 | Lecture 14: Generative Models 2

Unknown AuthorAbout 7 min readSep 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Generative vs. Discriminative Models: Different probabilistic models based on what is being predicted and normalized over.
  • Explicit vs. Implicit Density Models: Explicit models output a density value p(x), while implicit models allow sampling but don't provide p(x).
  • Generative Adversarial Networks (GANs): Implicit density models that use a generator and discriminator network in an adversarial training process.
  • Minimax Game: The adversarial objective function in GANs where the discriminator tries to maximize the value function V, and the generator tries to minimize it.
  • Diffusion Models: Generative models that convert noise into data through an iterative denoising process.
  • Rectified Flow Models: A specific type of diffusion model where noise is added through linear interpolation between data and noise samples.
  • Classifier Free Guidance (CFG): A technique to improve conditional generative models by training with and without conditioning information and combining the results.
  • Latent Diffusion Models: Diffusion models that operate in a latent space learned by an encoder-decoder network, often a VAE.
  • Distillation: Techniques to reduce the number of steps required for inference in diffusion models.
  • Score Function: The derivative of the log probability density function, used in some diffusion model formalisms.

Generative Adversarial Networks (GANs)

GANs vs. Autoregressive Models and VAEs

  • Autoregressive models and Variational Autoencoders (VAEs) are likelihood-based methods that directly model or approximate p(x) and maximize it.
  • GANs, on the other hand, give up on directly modeling p(x) but provide a way to sample from the underlying distribution.

GAN Setup and Training

  • Goal: To draw samples from a true data distribution pdata.
  • Latent Variable: Introduce a latent variable z, distributed according to a known prior p(z) (e.g., unit Gaussian).
  • Generator Network (G): Maps z to a generated data sample x from a generator distribution PG.
  • Discriminator Network (D): Classifies whether an image is real (from pdata) or fake (from PG).
  • Adversarial Training: Train G to fool D, and train D to correctly classify real vs. fake data.
  • Feedback: The generator receives feedback from the discriminator through gradients backpropagated through the generated image.

Minimax Game Equation

  • The generator (G) and discriminator (D) are jointly trained with a minimax game:
    • min_G max_D V(D, G) = E_{x~pdata(x)}[log D(x)] + E_{z~pz(z)}[log(1 - D(G(z)))]
  • Discriminator's Perspective:
    • Maximizes log D(x) for real data (D(x) close to 1).
    • Maximizes log(1 - D(G(z))) for fake data (D(G(z)) close to 0).
  • Generator's Perspective:
    • Minimizes log(1 - D(G(z))) to fool the discriminator (D(G(z)) close to 1).

Training Algorithm

  • Iterate:
    1. Update D: Gradient ascent on V with respect to D's parameters.
    2. Update G: Gradient descent on V with respect to G's parameters.

Challenges in GAN Training

  • V is not a loss function: The absolute value of V doesn't indicate how well PG matches pdata.
  • Unstable Objective: Jointly maximizing and minimizing the same quantity is difficult.
  • No Reliable Metric: Hard to tell when the model is doing a good job.

Training Dynamics and the Generator's Loss

  • Initial Stage: The generator produces random noise, making it easy for the discriminator to classify.
  • Low Gradients: The generator's loss function is flat at the beginning of training, making it hard to learn.
  • Hack: Instead of maximizing log(1 - D(G(z))), minimize -log(D(G(z))) for better gradients.

Theoretical Justification

  • The optimal discriminator can be written down mathematically (though not computed in practice).
  • The outer objective is minimized if and only if PG(x) = pdata(x).
  • Caveats: Assumes infinite capacity for G and D, and doesn't guarantee convergence.

GAN Architectures

  • DC-GAN: An early successful GAN with a five-layer ConvNet architecture.
  • StyleGAN: A more complex architecture that achieves high-quality results and smooth latent space interpolations.

Latent Space Interpolation

  • GANs tend to learn a smooth latent space, allowing for smooth transitions between generated samples when interpolating between latent vectors.

Pros and Cons of GANs

  • Pros: Simple formulation, can produce very nice results with careful tuning.
  • Cons: Unstable to train, no reliable loss curve, prone to mode collapse, hard to scale to big models and data.

Diffusion Models

Intuition Behind Diffusion Models

  • Goal: Convert samples from a noise distribution z into a data distribution px.
  • Constraint: The noise distribution z must have the same shape as the data x.
  • Noisy Data: Create versions of the data corrupted by increasing levels of noise (t from 0 to 1).
  • Denoising Network: Train a neural network to incrementally remove noise from a noisy sample.
  • Iterative Inference: Start with a noise sample and iteratively apply the neural network to remove noise.

Rectified Flow Models

  • Geometric Intuition:
    • Sample z from pnoise and x from pdata.
    • Draw a line (vector v) from x to z.
    • Set xt to be a point along this line (linear interpolation between x and z).
  • Training Objective: Train a neural network fθ to predict the vector v given xt and t.
  • Training Loop:
    1. Sample z from a unit Gaussian.
    2. Choose a noise level t uniformly from 0 to 1.
    3. Compute xt as a linear interpolation between x and z.
    4. Compute the loss as the mean squared error between the ground truth v and the model prediction.
  • Inference:
    1. Choose a number of steps T (e.g., 50).
    2. Sample x from the noise distribution.
    3. Iterate from t = 1 to 0:
      • Evaluate the model to get a predicted v.
      • Take a step along the predicted v to get a new x.

Conditional Rectified Flow

  • Setup: Data set has pairs (x, y), where y is a conditioning signal.
  • Model Input: The model takes y as an additional input.
  • Sampling: The model uses y during the iterative denoising process.

Classifier Free Guidance (CFG)

  • Training:
    • Train a conditional diffusion model that inputs xt and y.
    • With 50% probability, delete the conditioning information (set y to a null value).
  • Inference:
    • Evaluate the model twice: once with the conditioning signal (vy) and once without (v0).
    • Take a linear combination of the two vectors: v_CFG = vy + w * (vy - v0).
    • Step according to v_CFG.
  • Benefits: Allows controlling the strength of the conditioning signal.

Noise Schedules

  • Uniform Sampling: Puts uniform emphasis on all noise levels.
  • Logit-Normal Sampling: Puts more weight on intermediate noise levels.
  • Shifted Noise Schedules: Asymmetric schedules that shift more towards one direction or the other.

Latent Diffusion Models

  • Multi-Stage Procedure:
    1. Train an encoder network to map from image space to a latent space.
    2. Train a diffusion model on the latent space.
    3. At inference time, sample a random latent, denoise it with the diffusion model, and decode it with the decoder network.
  • Encoder-Decoder: Often a variational autoencoder (VAE) combined with a GAN to improve sample quality.

Neural Network Architectures

  • Diffusion Transformers (DiTs): Standard transformer blocks applied to diffusion models.
  • Conditioning Information: Injected through scale and shift mechanisms or sequence concatenation.

Applications

  • Text-to-Image Generation: Input a text prompt and generate an image.
  • Text-to-Video Generation: Input a text prompt and generate a video.

Challenges and Solutions

  • Slow Sampling: Distillation algorithms can reduce the number of steps required for inference.

Mathematical Formalisms

  • Latent Variable Model: Diffusion as maximizing a variational lower bound of the likelihood of the data.
  • Score Function Modeling: Diffusion as learning the score function of the data distribution.
  • Stochastic Differential Equations: Diffusion as solving a differential equation to transport samples from noise to data.

Conclusion

The lecture provided a detailed overview of generative models, focusing on GANs and diffusion models. GANs offer a unique adversarial training approach but are notoriously difficult to train. Diffusion models, particularly rectified flow models, provide a more stable training process and high-quality samples, especially when combined with latent spaces and techniques like classifier-free guidance. The modern generative modeling pipeline often integrates elements from VAEs, GANs, and diffusion models to achieve state-of-the-art results in tasks like text-to-image and text-to-video generation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.