Stanford CS231N Deep Learning for Computer Vision | Spring 2025 | Lecture 13: Generative Models 1

Unknown AuthorAbout 7 min readSep 3, 2025Watch original
THE SUMMARYAI-generated

CS231N Lecture 13: Generative Models - Summary

Key Concepts:

  • Self-Supervised Learning: Learning structure from unlabeled data using pretext tasks.
  • Contrastive Learning: Pulling similar pairs of data points together and pushing dissimilar pairs apart in feature space.
  • Generative Models: Models that learn the underlying probability distribution of data, allowing for the generation of new samples.
  • Discriminative Models: Models that learn to predict a label or target given an input.
  • Maximum Likelihood Estimation (MLE): A method for estimating the parameters of a probability distribution by maximizing the likelihood of observed data.
  • Autoregressive Models: Models that predict the next element in a sequence based on the previous elements.
  • Variational Autoencoders (VAEs): Probabilistic models that learn latent representations of data and can generate new samples by sampling from the latent space.
  • Evidence Lower Bound (ELBo): A lower bound on the log-likelihood of the data, used as the training objective for VAEs.

1. Self-Supervised Learning Recap

  • Two-Stage Procedure:
    1. Train an encoder-decoder on a pretext task using a large unlabeled dataset (e.g., billions of images from the internet).
    2. Discard the decoder, attach a small fully connected network to the encoder, and fine-tune on a small labeled dataset (e.g., tens, hundreds, or thousands of examples).
  • Pretext Tasks: Examples include rotation prediction, rearrangement (jigsaw puzzles), and reconstruction (inpainting). These tasks involve geometric perturbations to the input.
  • Contrastive Learning:
    • Uses similar and dissimilar pairs of data points.
    • Applies two random transformations to each input image (e.g., cropping, color jittering).
    • Feeds the transformed images to a feature extractor (e.g., ViT, CNN).
    • Computes a similarity matrix and pulls together feature vectors from augmentations of the same original image while pushing apart feature vectors from augmentations of different images.
  • SimCLR: A successful contrastive learning approach for self-supervised representation learning on images. Requires large batch sizes for good convergence.
  • MoCo (Momentum Contrast):
    • Addresses the large batch size requirement of SimCLR.
    • Maintains a queue (q) of samples from previous training iterations.
    • Uses a momentum encoder, which is an exponential moving average of the weights of the normal encoder, to process the queue.
    • The momentum encoder is not updated via gradient descent but through an exponential moving average of the normal encoder's weights.
  • DINO:
    • Similar to MoCo, using a dual encoder setup (normal and momentum).
    • Uses a KL divergence loss instead of a softmax.
  • DINO V2:
    • Scaled up DINO V1 to a much larger training set (142 million images vs. 1 million images in ImageNet).
    • Provides very strong self-supervised features and is widely used in practice.

2. Introduction to Generative Models

  • Generative models have seen significant progress in recent years, enabling tasks like language modeling, image generation, and video generation.
  • The fundamental ideas behind generative modeling have remained relatively consistent, but progress has been driven by increased compute, stable training recipes, larger datasets, and distributed training.

3. Supervised vs. Unsupervised Learning

  • Supervised Learning: Learning a function that maps inputs (x) to targets/labels (y) using a dataset of (x, y) pairs. Examples include image classification, image captioning, object detection, and segmentation.
  • Unsupervised Learning: Learning structure from data (x) without labels. Examples include K-means clustering, dimensionality reduction (PCA), and density estimation. The task itself is often unspecified.

4. Generative vs. Discriminative Models

  • Discriminative Models: Learn the conditional probability distribution p(y|x), i.e., the probability of a label given an input.
    • For every input x, the model predicts a probability distribution over all possible labels.
    • There is no competition among images for probability mass; only the labels for each image compete.
    • Have no real way to reject unreasonable inputs; they are forced to output a distribution over the fixed vocabulary.
  • Generative Models: Learn the joint probability distribution p(x), i.e., the probability of all possible images.
    • All possible images compete with each other for probability mass.
    • The model can reject unreasonable inputs by assigning low or zero probability mass.
  • Conditional Generative Models: Learn the conditional probability distribution p(x|y), i.e., the probability of an image given a label.
    • For every possible label, the model induces a competition among all possible images.
  • Bayes' Rule: Connects discriminative and generative models, allowing for the derivation of one from the others (in theory).

5. Use Cases for Different Model Types

  • Discriminative Models: Assign labels to data, feature learning.
  • Unconditional Generative Models: Detect outliers, feature learning without labels, sample and produce new samples (less practical due to lack of control).
  • Conditional Generative Models: Assign labels while rejecting outliers (less common), sample to generate new data from labels (most useful and interesting).

6. Why Generative Models?

  • Generative models are useful when there is ambiguity in the task being modeled.
  • They model a whole distribution of outputs conditioned on the input signal.
  • Examples:
    • Language Modeling: Predicting output text from input text (e.g., ChatGPT).
    • Text to Image: Generating images from text descriptions.
    • Image to Video: Predicting what happens next given an input image.

7. Taxonomy of Generative Models

  • Explicit Density Methods: Models where the probability density p(x) can be explicitly computed.
    • Direct: Can compute the real p(x). Example: Autoregressive Models.
    • Approximate: Can compute an approximation to the true density. Example: Variational Autoencoders (VAEs).
  • Implicit Density Methods: Models where the probability density p(x) cannot be explicitly computed, but samples can be drawn from the distribution.
    • Direct: Requires a single network evaluation to draw a sample. Example: Generative Adversarial Networks (GANs).
    • Indirect: Requires an iterative procedure to draw a sample. Example: Diffusion Models.

8. Autoregressive Models

  • Maximum Likelihood Estimation (MLE): A general procedure for fitting probabilistic models by maximizing the likelihood of the data.
    • Assumes an underlying true probability distribution p_data.
    • Maximizes the log-likelihood of the data.
  • Autoregressive Assumption: Data can be split into a sequence of subparts (x1, x2, ..., xT).
  • Chain Rule of Probability: p(x) = p(x1) * p(x2|x1) * p(x3|x1, x2) * ... * p(xT|x1, ..., xT-1).
  • RNNs and Transformers: Naturally suited for autoregressive modeling due to their sequential processing capabilities.
  • Application to Text: Text data is naturally a 1D sequence of discrete tokens, making it well-suited for autoregressive models.
  • Application to Images:
    • Images can be treated as a sequence of pixels.
    • Each pixel is composed of three discrete values (RGB).
    • Rasterizing the image into a long sequence of pixel values allows for autoregressive modeling.
    • This approach can be computationally expensive for high-resolution images.

9. Variational Autoencoders (VAEs)

  • Motivation: To learn latent representations (z) of data (x) and generate new samples by sampling from the latent space.
  • Autoencoders (AEs):
    • Unsupervised method for learning to extract features from inputs without labels.
    • Consist of an encoder and a decoder.
    • The encoder maps the input (x) to a latent representation (z).
    • The decoder maps the latent representation (z) back to the input (x).
    • Trained to reconstruct the input, forcing the model to learn useful features by bottlenecking the representation (z).
  • Variational Autoencoders (VAEs):
    • A probabilistic spin on traditional autoencoders.
    • Enforce a structure on the latent space (z) to enable sampling.
    • Assume that each training data point (xi) was generated from an underlying latent vector (zi).
    • Force the latent vectors (z) to come from a known distribution (typically a unit Gaussian).
  • Training Objective: Maximize the Evidence Lower Bound (ELBo) on the log-likelihood of the data.
    • The ELBo consists of a data reconstruction term and a prior term.
    • The data reconstruction term encourages the model to reconstruct the input from the latent representation.
    • The prior term encourages the latent space to match the assumed prior distribution (e.g., unit Gaussian).
  • Reparameterization Trick: Allows for backpropagation through the sampling process.
  • Loss Function: The two losses (reconstruction and prior) fight against each other, forcing the model to find a balance between reconstructing the data well and maintaining a latent space close to the prior.
  • Sampling: After training, samples can be generated by sampling from the prior distribution and passing them through the decoder.
  • Statistical Independence: The diagonal Gaussian latent space provides a notion of statistical independence across the different entries in the latent space, allowing for interpretable variations.

10. Conclusion

The lecture provided a comprehensive overview of generative models, covering the theoretical foundations, different types of models, and practical considerations for training and using them. It highlighted the importance of generative models for tasks involving ambiguity and the potential for generating new and diverse data samples. The lecture also delved into the specifics of autoregressive models and variational autoencoders, explaining their architectures, training objectives, and key concepts. The next lecture will cover the other half of the generative model family tree, focusing on GANs and diffusion models.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.