Stanford CS231N | Spring 2025 | Lecture 12: Self-Supervised Learning

Unknown AuthorAbout 7 min readSep 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Self-Supervised Learning (SSL): Training neural networks without manual labels, using pretext tasks to learn useful features.
  • Pretext Task: A task designed to generate labels automatically from unlabeled data, enabling feature learning. Examples include image rotation prediction, jigsaw puzzle solving, image inpainting, and colorization.
  • Downstream Task: The actual task of interest (e.g., classification, object detection) that uses the features learned during pre-training.
  • Encoder: The part of the neural network that learns to extract meaningful features from the input data.
  • Decoder: The part of the neural network that maps the learned representations (features) to the output space, often used in pretext tasks.
  • Learned Representations/Features/Embeddings/Latent Space: Different names for the output of the encoder, representing the learned characteristics of the input data.
  • Contrastive Learning: A self-supervised learning approach that aims to bring representations of similar data points closer together and push representations of dissimilar data points further apart in the feature space.
  • Masked Autoencoders (MAE): A reconstruction-based self-supervised learning framework that masks a large portion of the input data and trains the model to reconstruct the missing parts.
  • InfoNCE (Information Noise Contrastive Estimation): A loss function used in contrastive learning to maximize the mutual information between positive pairs and minimize it between negative pairs.
  • SimCLR (Simple Framework for Contrastive Learning): A contrastive learning framework that uses data augmentations to create positive pairs and maximizes the similarity between their representations.
  • MoCo (Momentum Contrastive Learning): A contrastive learning framework that uses a momentum encoder and a queue of negative samples to improve training stability and performance.
  • DINO: A self-supervised learning framework that uses a teacher-student network architecture and does not necessarily rely on contrastive learning.

Self-Supervised Learning: Bypassing Manual Labels

The lecture focuses on self-supervised learning (SSL) as a method to train neural networks without relying on large, manually labeled datasets. The core idea is to define a "pretext task" that allows the network to learn useful features from unlabeled data. These features can then be transferred to a "downstream task" of interest, where only a small amount of labeled data is available.

  • The Challenge of Labeled Data: Training large-scale networks requires vast amounts of labeled data, which is expensive and time-consuming to acquire, especially for tasks like segmentation where pixel-level annotation is needed.
  • The Self-Supervised Learning Paradigm: SSL aims to bypass the need for manual labels by defining a pretext task that generates its own labels from the data itself.
  • Pretext Task and Downstream Task: The process involves pre-training an encoder using a pretext task on a large unlabeled dataset. The trained encoder is then transferred to a downstream task, where it is fine-tuned with a small labeled dataset.
  • Encoder and Decoder: The neural network often consists of an encoder that extracts features and a decoder, classifier, or regressor that maps those features to the output space of the pretext task.

Pretext Tasks: Learning from Unlabeled Data

The lecture explores several examples of pretext tasks used in self-supervised learning:

1. Image Rotation Prediction

  • Task: Rotate an image by a specific angle (e.g., 0, 90, 180, 270 degrees) and train the network to predict the rotation angle.
  • Rationale: The model must develop a "visual common sense" of how objects should look in their correct orientation to accurately predict the rotation.
  • Implementation: A convolutional neural network (CNN) is used to classify the rotation angle.
  • Evaluation: The learned representations are evaluated by fine-tuning the network on downstream tasks like image classification, object detection, and segmentation.
  • Results: Pre-training with rotation prediction significantly improves performance compared to random initialization, approaching the performance of supervised pre-training on ImageNet.
  • Feature Visualization: Attention maps from self-supervised models tend to cover more areas of the image compared to supervised models, indicating a more holistic understanding.

2. Jigsaw Puzzle Solving

  • Task: Divide an image into patches, shuffle them randomly, and train the network to predict the correct permutation of the patches.
  • Implementation: The image is divided into a 3x3 grid, and the network is trained to predict the correct permutation out of a set of 64 plausible permutations.
  • Rationale: The model must understand the spatial relationships between different parts of the image to solve the puzzle.

3. Image Inpainting (Completion)

  • Task: Mask parts of an image and train the network to reconstruct the missing pixels.
  • Implementation: A masking strategy is used to hide parts of the input image, and the network is trained to reconstruct the missing regions.
  • Autoencoder Framework: This task is often implemented using an autoencoder architecture, where the encoder maps the masked image to a feature space, and the decoder reconstructs the missing pixels.
  • Adversarial Objective: To generate more realistic and less blurry reconstructions, an adversarial objective function is often added to the loss function.
  • Loss Function: The loss function typically includes a reconstruction loss (e.g., mean squared error) calculated only on the masked area.

4. Image Colorization

  • Task: Given a black and white version of an image, train the network to predict the colors for each pixel.
  • Color Spaces: The image is converted to a color space that separates lightness (illumination) from color, such as the LAB color space.
  • Split-Brain Autoencoder: A variation of this task involves splitting the image into lightness and color channels and training two separate networks to predict the other channel.
  • Applications: Image colorization can be used to colorize old photos and videos.
  • Video Colorization: By colorizing a reference frame and then tracking pixels and objects in subsequent frames, the model can learn to track regions and objects without labels.

Masked Autoencoders (MAE): A Modern Approach

  • Framework: MAE is a reconstruction-based self-supervised learning framework that masks a large portion of the input data (e.g., 75%) and trains the model to reconstruct the missing parts.
  • Encoder and Decoder: The encoder processes the unmasked patches, and the decoder reconstructs the entire image, including the masked regions.
  • Transformer Architecture: Both the encoder and decoder are typically based on transformer architectures, similar to ViTs (Vision Transformers).
  • Masking Strategy: Random masking is used to select the patches to be masked.
  • High Masking Ratio: Using a high masking ratio (e.g., 75%) makes the prediction task more challenging and forces the model to learn more meaningful features.
  • Training: The model is trained using a mean squared error (MSE) loss function calculated only on the masked patches.
  • Fine-tuning: The pre-trained encoder can be fine-tuned on downstream tasks using either linear probing (freezing the encoder and training a linear classifier) or full fine-tuning (fine-tuning the entire network).
  • Performance: MAE has shown state-of-the-art performance on various downstream tasks, outperforming other self-supervised learning methods.

Contrastive Learning: Learning by Comparison

  • Core Idea: Contrastive learning aims to learn representations by bringing representations of similar data points closer together and pushing representations of dissimilar data points further apart in the feature space.
  • Positive and Negative Samples: The process involves defining positive samples (e.g., different transformations of the same image) and negative samples (e.g., images from different objects).
  • Scoring Function: A scoring function is used to measure the similarity between representations.
  • InfoNCE Loss: The InfoNCE (Information Noise Contrastive Estimation) loss function is used to maximize the similarity between positive pairs and minimize the similarity between negative pairs.
  • Batch Learning: Contrastive learning is often implemented using batch learning, where all other objects in the batch are considered as negative samples.
  • SimCLR: SimCLR is a contrastive learning framework that uses data augmentations to create positive pairs and maximizes the similarity between their representations.
  • MoCo: MoCo is a contrastive learning framework that uses a momentum encoder and a queue of negative samples to improve training stability and performance.

Key Takeaways

  • Self-supervised learning offers a promising approach to training neural networks without relying on large, manually labeled datasets.
  • Pretext tasks are crucial for learning useful features from unlabeled data.
  • Masked Autoencoders (MAE) have emerged as a powerful framework for self-supervised learning.
  • Contrastive learning provides an alternative approach to learning representations by comparing similar and dissimilar data points.
  • The choice of pretext task, network architecture, and training parameters can significantly impact the performance of self-supervised learning models.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.