Stanford CS231N | Spring 2025 | Lecture 7: Recurrent Neural Networks

Unknown AuthorAbout 9 min readSep 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Dropout Scaling: Adjusting probabilities at test time to match expected output during training.
  • Normalization (LayerNorm): A technique to improve training stability, but may not always be beneficial depending on the problem.
  • Vanilla RNN: A basic recurrent neural network with a fixed-size input and output, using a recurrence formula and activation functions.
  • Sequence Modeling: Processing variable-length sequences of data, including one-to-many, many-to-one, and many-to-many scenarios.
  • Hidden State: An internal state in RNNs that is updated as a sequence is processed, capturing information from previous time steps.
  • Unrolled RNN: A diagram representing the explicit dependencies in an RNN, showing how the current hidden state depends on the input and previous hidden state.
  • Backpropagation Through Time (BPTT): Computing gradients in RNNs by summing the gradients at each time step.
  • Truncated BPTT: A technique to reduce memory usage during training by fixing a time window and batching the computational graph.
  • Character-Level Language Model: An RNN that predicts the next character in a sequence, trained on a large corpus of text.
  • Embedding Layer: A matrix that maps input indices to dense vectors, improving model performance compared to one-hot encodings.
  • Greedy Decoding: Always picking the maximum probability output at each time step, which can lead to repetitive sequences.
  • Sampling: Choosing outputs based on a distribution derived from the softmax probabilities, introducing variability in generated sequences.
  • LSTM (Long Short-Term Memory): A variant of RNNs with gates to control information flow, addressing the vanishing gradient problem.
  • State Space Models (SSMs): Modern language models inspired by RNNs, aiming for linear time sequence modeling.
  • Context Length: The maximum length of input sequence a model can process, a limitation of transformers that RNNs overcome.

Clarifications and Initial Remarks

  • Dropout Scaling at Test Time: The hyperparameter p in dropout represents the probability of dropping out neurons (or keeping them active, depending on implementation). At test time, activations should be scaled by (1-p) to maintain the same expected output as during training. The original slide had a mismatch in the definition of p.
  • Normalization and Weight Initialization: LayerNorm can help with issues arising from incorrect weight initialization, but good initialization (e.g., Kaiming initialization) is still crucial for optimal performance. The effectiveness of LayerNorm depends on the problem; it can hurt performance if precise spatial information is needed.
  • Recap of Vanilla Neural Networks: The lecture transitions from discussing clarifications to sequence modeling, recapping key aspects of vanilla neural networks: fixed-size input/output, activation functions, data pre-processing, weight initialization/normalization, transfer learning, training dynamics (learning rate, hyperparameter optimization), and test-time augmentation.
  • Weights & Biases Tool: Weights & Biases is recommended as a tool for visualizing and optimizing hyperparameters based on validation set performance. It allows for easy comparison of different runs with varying hyperparameters.

Sequence Modeling and RNNs

  • Sequence Modeling Task: The lecture introduces sequence modeling, contrasting it with fixed-size input/output models. Different sequence modeling scenarios are presented:
    • One-to-many: Fixed-size input to variable-length output (e.g., image captioning).
    • Many-to-one: Variable-length input to fixed-size output (e.g., video classification).
    • Many-to-many: Variable-length input to variable-length output (e.g., video captioning, frame-by-frame video classification).
  • RNN Architecture: RNNs process input sequences x and produce output sequences y. The key feature is the recurrent nature, where an internal (hidden) state is updated as the sequence is processed. The current hidden state depends on the current input and the previous hidden state.
  • Mathematical Formulation: The new hidden state h_t is a function of the old hidden state h_{t-1} and the input vector x_t at time step t, using a function f with parameters W (typically a weight matrix and an activation function). The output y_t is calculated using a separate function g with parameters W_hy, transforming the hidden state to the output dimension.
  • Initialization: The initial hidden state h_0 needs to be initialized. It can be a learned vector or set to a fixed value.
  • Vanilla RNN Details: A Vanilla RNN uses the hyperbolic tangent (tanh) as the activation function, bounded between -1 and 1, zero-centered, and capable of representing both positive and negative values. The output y_t can be a matrix multiplied by the hidden state h_t.

Concrete Example: Detecting Repeated Ones

  • Task Definition: The lecture presents a toy example: creating an RNN that outputs 1 when there are two consecutive 1s in an input sequence of 0s and 1s, and 0 otherwise.
  • Hidden State Design: The hidden state h_t is a 3-dimensional vector: [current value, previous value, 1]. It is initialized to [0, 0, 1], representing two initial zeros.
  • Weight Matrices: Two weight matrices are defined:
    • W_xh: Converts the input x to the hidden state dimension. Set to [1, 0, 0]^T, so that when x is 0, we get a 0 vector, and when x is 1, we get [1,0,0]^T.
    • W_hh: Transforms the previous hidden state to the next one. The top row is all zeros, so that the current value is not changed based on the previous hidden state. The second row is [1, 0, 0], copying the current value from the previous time step to the previous value of the current time step. The third row is [0,0,1] to maintain the 1.
  • Output Calculation: The output y_t is calculated as ReLU(current + previous - 1). This formula outputs 1 only when both the current and previous values are 1.
  • ReLU Activation: ReLU (Rectified Linear Unit) is used as the activation function (max(0, value)) to simplify the math.

Computing Gradients and Backpropagation Through Time

  • Computational Graph: The lecture explains how to compute gradients in RNNs using the computational graph. The same weight matrices W are used at each time step.
  • Loss Calculation: In a many-to-many scenario, a loss can be calculated for each output, and the total loss is the sum of these individual losses. Gradients can be calculated for each time step separately and then summed together.
  • Many-to-One Scenario: In a many-to-one scenario, a single loss is calculated, often using only the final hidden state.
  • One-to-Many Scenario: In a one-to-many scenario (e.g., image captioning), the previous outputs y can be incorporated as input to the next time step.
  • Backpropagation Through Time (BPTT) Issues: With long input sequences, storing activations and gradients at each time step can lead to memory issues.
  • Truncated BPTT Solution: Truncated BPTT fixes a time window and treats each window as a separate training example. The hidden state is initialized with the output of the previous window, but gradients are not carried over. This reduces memory usage.
  • Gradient Calculation with Chunking: When using truncated BPTT, the gradient of the loss with respect to the final hidden state of a chunk is calculated. This gradient is then used to calculate the gradients for the previous time steps within the chunk.

Character-Level Language Models

  • Model Description: Character-level language models predict the next character in a sequence. Characters are typically represented using one-hot encoding.
  • Training Process: The model is trained to minimize the loss between the predicted character and the actual next character.
  • Test Time: At test time, the model samples characters one at a time and feeds them back as input to generate text.
  • Embedding Layer: Instead of one-hot encodings, an embedding layer is often used. This layer maps input indices to dense vectors, which are learned during training.
  • Example Results: The lecture shows examples of character-level RNNs trained on sonnets by William Shakespeare and Linux source code, demonstrating their ability to generate coherent text and code.
  • Modern Coding Tools: Modern programming tools based on language models use a similar approach, predicting the next token (group of characters) instead of individual characters.
  • Sampling from the Model: Instead of always picking the maximum probability output (greedy decoding), the model samples from a distribution based on the softmax probabilities. This introduces variability in the generated sequences. Beam search is also mentioned as a technique to search ahead and find the sequence with the highest overall probability.

Analyzing RNN Activations

  • Activation Visualization: The lecture discusses visualizing RNN activations to understand what the model is tracking. Activation values are mapped to colors, and these colors are plotted for each character in the input sequence.
  • Interpretable Cells: Some cells in the RNN learn to track specific features, such as quote detection, line length, if statements, code depth, and comments.

Advantages and Disadvantages of RNNs

  • Advantages:
    • Can process any length of input.
    • Can theoretically use information from many steps back.
    • Model size does not increase for longer inputs.
    • Same weights are applied at each time step.
  • Disadvantages:
    • Need to compute the previous hidden state to compute the next one, which can be slow.
    • Difficult to batch all time steps together during training.
    • Difficult to access information many time steps back due to the fixed-size hidden state.

Applications of RNNs in Computer Vision

  • Image Captioning: A CNN encodes the image, and an RNN generates a caption based on the image encoding and the previously generated text. The process terminates when an end token is generated.
  • Visual Question Answering: RNNs can be used to answer questions about images. Two common formulations are:
    • Using a captioning model and calculating the probability of each answer sequence.
    • Inputting the question and multiple answers as separate inputs and outputting a probability for each answer.
  • Visual Dialogue: RNNs can be used to have a chat about an image.
  • Visual Navigation: RNNs can be used to output a sequence of directions to move in a 2D floor plan to reach a target destination.

Multi-Layer RNNs

  • Architecture: Multi-layer RNNs stack multiple RNN layers on top of each other. The hidden state of each layer depends on the hidden state of the previous time step within that layer. The input to the second layer is the output of the first layer, and so on.
  • Computational Complexity: Training multi-layer RNNs can be computationally expensive.

LSTMs (Long Short-Term Memory)

  • Motivation: LSTMs were developed to address the vanishing gradient problem and the difficulty of capturing long-term dependencies in RNNs.
  • Gating Mechanism: LSTMs use gates to control the flow of information:
    • Input gate: Decides whether to write information to the cell.
    • Forget gate: Decides how much to forget from previous time steps.
    • Output gate: Decides how much to output for the hidden state.
  • Highway: LSTMs have a separate pathway (highway) where information can be passed more easily without activation functions, helping to preserve long-term dependencies.
  • ResNet Connection: The skip connections in ResNets are related to the highway in LSTMs, both helping to preserve information over many layers or time steps.

Resurgence of RNNs and State Space Models

  • Advantages of RNNs:
    • Unlimited context length.
    • Linear compute scaling with sequence length.
  • State Space Models (SSMs): Modern language models inspired by RNNs, such as RWKV and Mamba, aim for linear time sequence modeling.
  • Goal: To achieve the performance of transformers with the scaling of RNNs.

Conclusion

The lecture concludes by summarizing the key points: Vanilla RNNs are simple but have limitations. More complex variants like LSTMs introduce ways to selectively pass information. The backward flow of gradients in RNNs can either explode or vanish. Better architectures and new paradigms for reasoning over sequences are hot topics of research. The lecture sets the stage for the next lecture on attention and transformers.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.