Key Concepts
- Generative vs. Discriminative Models: Different probabilistic models based on what is being predicted and normalized over.
- Explicit vs. Implicit Density Models: Explicit models output a density value p(x), while implicit models allow sampling but don't provide p(x).
- Generative Adversarial Networks (GANs): Implicit density models that use a generator and discriminator network in an adversarial training process.
- Minimax Game: The adversarial objective function in GANs where the discriminator tries to maximize the value function V, and the generator tries to minimize it.
- Diffusion Models: Generative models that convert noise into data through an iterative denoising process.
- Rectified Flow Models: A specific type of diffusion model where noise is added through linear interpolation between data and noise samples.
- Classifier Free Guidance (CFG): A technique to improve conditional generative models by training with and without conditioning information and combining the results.
- Latent Diffusion Models: Diffusion models that operate in a latent space learned by an encoder-decoder network, often a VAE.
- Distillation: Techniques to reduce the number of steps required for inference in diffusion models.
- Score Function: The derivative of the log probability density function, used in some diffusion model formalisms.
Generative Adversarial Networks (GANs)
GANs vs. Autoregressive Models and VAEs
- Autoregressive models and Variational Autoencoders (VAEs) are likelihood-based methods that directly model or approximate p(x) and maximize it.
- GANs, on the other hand, give up on directly modeling p(x) but provide a way to sample from the underlying distribution.
GAN Setup and Training
- Goal: To draw samples from a true data distribution pdata.
- Latent Variable: Introduce a latent variable z, distributed according to a known prior p(z) (e.g., unit Gaussian).
- Generator Network (G): Maps z to a generated data sample x from a generator distribution PG.
- Discriminator Network (D): Classifies whether an image is real (from pdata) or fake (from PG).
- Adversarial Training: Train G to fool D, and train D to correctly classify real vs. fake data.
- Feedback: The generator receives feedback from the discriminator through gradients backpropagated through the generated image.
Minimax Game Equation
- The generator (G) and discriminator (D) are jointly trained with a minimax game:
min_G max_D V(D, G) = E_{x~pdata(x)}[log D(x)] + E_{z~pz(z)}[log(1 - D(G(z)))]
- Discriminator's Perspective:
- Maximizes
log D(x)for real data (D(x) close to 1). - Maximizes
log(1 - D(G(z)))for fake data (D(G(z)) close to 0).
- Maximizes
- Generator's Perspective:
- Minimizes
log(1 - D(G(z)))to fool the discriminator (D(G(z)) close to 1).
- Minimizes
Training Algorithm
- Iterate:
- Update D: Gradient ascent on V with respect to D's parameters.
- Update G: Gradient descent on V with respect to G's parameters.
Challenges in GAN Training
- V is not a loss function: The absolute value of V doesn't indicate how well PG matches pdata.
- Unstable Objective: Jointly maximizing and minimizing the same quantity is difficult.
- No Reliable Metric: Hard to tell when the model is doing a good job.
Training Dynamics and the Generator's Loss
- Initial Stage: The generator produces random noise, making it easy for the discriminator to classify.
- Low Gradients: The generator's loss function is flat at the beginning of training, making it hard to learn.
- Hack: Instead of maximizing
log(1 - D(G(z))), minimize-log(D(G(z)))for better gradients.
Theoretical Justification
- The optimal discriminator can be written down mathematically (though not computed in practice).
- The outer objective is minimized if and only if PG(x) = pdata(x).
- Caveats: Assumes infinite capacity for G and D, and doesn't guarantee convergence.
GAN Architectures
- DC-GAN: An early successful GAN with a five-layer ConvNet architecture.
- StyleGAN: A more complex architecture that achieves high-quality results and smooth latent space interpolations.
Latent Space Interpolation
- GANs tend to learn a smooth latent space, allowing for smooth transitions between generated samples when interpolating between latent vectors.
Pros and Cons of GANs
- Pros: Simple formulation, can produce very nice results with careful tuning.
- Cons: Unstable to train, no reliable loss curve, prone to mode collapse, hard to scale to big models and data.
Diffusion Models
Intuition Behind Diffusion Models
- Goal: Convert samples from a noise distribution z into a data distribution px.
- Constraint: The noise distribution z must have the same shape as the data x.
- Noisy Data: Create versions of the data corrupted by increasing levels of noise (t from 0 to 1).
- Denoising Network: Train a neural network to incrementally remove noise from a noisy sample.
- Iterative Inference: Start with a noise sample and iteratively apply the neural network to remove noise.
Rectified Flow Models
- Geometric Intuition:
- Sample z from pnoise and x from pdata.
- Draw a line (vector v) from x to z.
- Set xt to be a point along this line (linear interpolation between x and z).
- Training Objective: Train a neural network fθ to predict the vector v given xt and t.
- Training Loop:
- Sample z from a unit Gaussian.
- Choose a noise level t uniformly from 0 to 1.
- Compute xt as a linear interpolation between x and z.
- Compute the loss as the mean squared error between the ground truth v and the model prediction.
- Inference:
- Choose a number of steps T (e.g., 50).
- Sample x from the noise distribution.
- Iterate from t = 1 to 0:
- Evaluate the model to get a predicted v.
- Take a step along the predicted v to get a new x.
Conditional Rectified Flow
- Setup: Data set has pairs (x, y), where y is a conditioning signal.
- Model Input: The model takes y as an additional input.
- Sampling: The model uses y during the iterative denoising process.
Classifier Free Guidance (CFG)
- Training:
- Train a conditional diffusion model that inputs xt and y.
- With 50% probability, delete the conditioning information (set y to a null value).
- Inference:
- Evaluate the model twice: once with the conditioning signal (vy) and once without (v0).
- Take a linear combination of the two vectors:
v_CFG = vy + w * (vy - v0). - Step according to v_CFG.
- Benefits: Allows controlling the strength of the conditioning signal.
Noise Schedules
- Uniform Sampling: Puts uniform emphasis on all noise levels.
- Logit-Normal Sampling: Puts more weight on intermediate noise levels.
- Shifted Noise Schedules: Asymmetric schedules that shift more towards one direction or the other.
Latent Diffusion Models
- Multi-Stage Procedure:
- Train an encoder network to map from image space to a latent space.
- Train a diffusion model on the latent space.
- At inference time, sample a random latent, denoise it with the diffusion model, and decode it with the decoder network.
- Encoder-Decoder: Often a variational autoencoder (VAE) combined with a GAN to improve sample quality.
Neural Network Architectures
- Diffusion Transformers (DiTs): Standard transformer blocks applied to diffusion models.
- Conditioning Information: Injected through scale and shift mechanisms or sequence concatenation.
Applications
- Text-to-Image Generation: Input a text prompt and generate an image.
- Text-to-Video Generation: Input a text prompt and generate a video.
Challenges and Solutions
- Slow Sampling: Distillation algorithms can reduce the number of steps required for inference.
Mathematical Formalisms
- Latent Variable Model: Diffusion as maximizing a variational lower bound of the likelihood of the data.
- Score Function Modeling: Diffusion as learning the score function of the data distribution.
- Stochastic Differential Equations: Diffusion as solving a differential equation to transport samples from noise to data.
Conclusion
The lecture provided a detailed overview of generative models, focusing on GANs and diffusion models. GANs offer a unique adversarial training approach but are notoriously difficult to train. Diffusion models, particularly rectified flow models, provide a more stable training process and high-quality samples, especially when combined with latent spaces and techniques like classifier-free guidance. The modern generative modeling pipeline often integrates elements from VAEs, GANs, and diffusion models to achieve state-of-the-art results in tasks like text-to-image and text-to-video generation.
AI summaries can miss context or contain errors. Check important details against the original video.





