Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 2 - Transformer-Based Models & Tricks
By Unknown Author
Key Concepts
- Self-Attention: A mechanism where each token in a sequence attends to all other tokens to compute its representation.
- Queries, Keys, Values: Components used in self-attention to determine token similarity and extract information.
- Transformer Architecture: A neural network architecture composed of an encoder and a decoder, initially designed for machine translation.
- Multi-Head Attention: An extension of self-attention where multiple attention mechanisms (heads) operate in parallel, allowing the model to learn different projection aspects.
- Attention Map: A visualization representing the attention weights between tokens, showing which tokens are most relevant to each other.
- Position Embedding: A method to inject positional information into token embeddings, as self-attention itself is permutation-invariant.
- Learned Position Embeddings: Position embeddings that are learned during training, with limitations in handling unseen sequence lengths.
- Sinusoidal Position Embeddings: A deterministic method for position embeddings using sine and cosine functions, allowing for arbitrary sequence lengths.
- Relative Position Bias: A technique to incorporate positional information directly into the attention mechanism by adding a bias term based on relative distances.
- ALiBi (Attention with Linear Bias): A deterministic approach to relative position bias using a linear formula.
- RoPE (Rotary Position Embeddings): A method that rotates query and key vectors by angles dependent on their positions, capturing relative distances.
- Layer Normalization (LayerNorm): A technique to stabilize training by normalizing the activations within a layer.
- Pre-Norm vs. Post-Norm: Variations in the placement of LayerNorm within the transformer architecture.
- RMSNorm (Root Mean Square Normalization): A simplified normalization technique that uses the root mean square of activations.
- Sliding Window Attention: A variation of attention that restricts interactions to a local window of tokens, reducing computational complexity.
- Multi-Query Attention (MQA): A method where multiple attention heads share the same projection matrices for keys and values.
- Group Query Attention (GQA): A compromise between MQA and standard multi-head attention, where projection matrices are shared across groups of heads.
- Encoder-Decoder Architecture: The original transformer structure with separate encoder and decoder stacks.
- Encoder-Only Architecture: Models that utilize only the encoder part of the transformer, suitable for classification and representation tasks.
- Decoder-Only Architecture: Models that utilize only the decoder part of the transformer, primarily for generative tasks.
- BERT (Bidirectional Encoder Representations from Transformers): An influential encoder-only model pre-trained using Masked Language Model (MLM) and Next Sentence Prediction (NSP) objectives.
- WordPiece Tokenizer: A subword tokenization algorithm used by BERT.
- CLS Token: A special token added to the beginning of an input sequence in BERT, used to aggregate the sequence's representation for classification.
- SEP Token: A special token used in BERT to separate sentences.
- Segment Encoding: An embedding added to tokens to indicate which sentence they belong to, used in BERT for NSP.
- Masked Language Model (MLM): A pre-training objective where some input tokens are masked, and the model must predict them based on their context.
- Next Sentence Prediction (NSP): A pre-training objective where the model predicts whether two sentences are consecutive in the original corpus.
- Pre-training: The initial stage of training a model on a large, unlabeled dataset to learn general representations.
- Fine-tuning: The subsequent stage where a pre-trained model is adapted to a specific downstream task using a smaller, labeled dataset.
- Distillation: A technique where a smaller "student" model learns to mimic the output distribution of a larger "teacher" model.
- KL Divergence: A measure of the difference between two probability distributions, used as a loss function in distillation.
- DistilBERT: A distilled version of BERT, smaller and faster while retaining much of BERT's performance.
- RoBERTa (Robustly Optimized BERT Approach): An optimized version of BERT that modifies pre-training objectives and data strategies.
Lecture 2: Transformer Architecture and Modern Variations
This lecture builds upon the concept of self-attention introduced in Lecture 1, focusing on key components of the transformer architecture that have evolved and their impact on modern models.
1. Recap of Lecture 1: Self-Attention and Transformer Basics
- Self-Attention: Each token attends to all other tokens in a sequence using queries, keys, and values. The similarity between a query and keys determines attention weights, which are then applied to values.
- Formula: The self-attention mechanism is expressed as
softmax(QK^T / sqrt(d_k)) * V, whereQ,K, andVare matrices representing queries, keys, and values, andd_kis the dimension of the keys. This formula is highly optimized for hardware. - Transformer Architecture: Consists of an encoder (processing input) and a decoder (generating output), initially for machine translation.
- Multi-Head Attention: Involves multiple "heads," each learning a different projection of the input into queries, keys, and values. This allows the model to capture diverse relationships.
- Attention Maps: Visualizations showing attention weights, helping to interpret which tokens are most attended to. For example, the token "its" might attend strongly to "law" and "application" in a sentence, indicating its referential context. Each head can learn different patterns of attention.
- Projection Matrices: Each attention head uses its own projection matrices for queries, keys, and values. Computations are parallelized, with results concatenated and projected again.
2. Key Evolving Components of the Transformer
The lecture focuses on three critical components of the transformer that have seen significant modifications: position embeddings, layer normalization, and the attention mechanism itself.
2.1. Position Embeddings
-
Problem: Self-attention is permutation-invariant, meaning it doesn't inherently understand the order of tokens. This positional information is crucial for language understanding.
-
Original Transformer Approach (Learned Embeddings):
- Each position in the sequence is assigned a unique, learnable embedding.
- This position embedding is added to the token embedding.
- Limitations:
- Data Dependency: Learned embeddings are biased by the positions present in the training data.
- Sequence Length Limit: Can only learn embeddings up to the maximum sequence length seen during training. Inference with longer sequences is problematic.
-
Sinusoidal Position Embeddings (Original Transformer Paper):
- A deterministic formula using sine and cosine functions to generate position embeddings.
- For a position
mand dimensioni(withind_model), the embedding is calculated usingsin(m * omega_i)andcos(m * omega_i), whereomega_i = 10000^(-2i / d_model). - Intuition: This formulation ensures that the dot product between position embeddings
mandnis a function of their relative distance (m - n). This mimics the intuition that closer tokens are more relevant. - Properties:
- Extensibility: Can handle arbitrary sequence lengths beyond the training set.
- Relative Distance Encoding: The dot product between embeddings of positions
mandnis related tocos(omega_i * (m - n)). - Frequency Variation: Lower dimensions vary rapidly (high frequency), while higher dimensions vary slowly (low frequency), allowing for encoding of different relative distances.
- Performance: Achieved comparable results to learned embeddings but with the advantage of handling longer sequences.
-
Modern Approaches to Position Information:
- Direct Intervention in Attention: Instead of adding to input embeddings, positional information is integrated directly into the attention calculation.
- Relative Position Bias (T5): Learns a bias term that is added to the attention scores based on the relative distance between tokens. This bias is often bucketized.
- ALiBi (Attention with Linear Bias): A deterministic approach that adds a linear bias based on the relative distance, avoiding learned parameters for this aspect.
- RoPE (Rotary Position Embeddings):
- Rotates query and key vectors by angles dependent on their positions.
- This rotation is achieved through matrix multiplication with a rotation matrix.
- Key Property: The dot product of rotated queries and keys becomes a function of the relative distance (
n - m). - Mathematical Basis: Extends to
d-dimensional space by treating pairs of dimensions as 2D planes for rotation. The anglethetais a function of the dimension indexiandd_model. - Impact: Widely adopted in modern LLMs due to its ability to capture relative positional information effectively and its favorable properties for attention scores (e.g., long-term decay).
2.2. Layer Normalization
- Purpose: To stabilize training and speed up convergence by normalizing the activations within a layer. It addresses the issue of "internal covariate shift," where activation distributions change significantly across layers and training steps.
- Original Transformer (Post-Norm):
- The input and the output of a sub-layer (attention or FFN) are added, and then LayerNorm is applied.
- Formula:
LayerNorm(x + Sublayer(x))
- Modern Transformers (Pre-Norm):
- LayerNorm is applied before the input enters the sub-layer.
- Formula:
x + Sublayer(LayerNorm(x)) - Benefit: Empirically shown to lead to more stable training and faster convergence.
- RMSNorm:
- A simplified normalization technique that normalizes by the root mean square of the activations.
- It learns only a rescaling factor (
gamma), reducing the number of learnable parameters compared to LayerNorm. - Offers comparable convergence properties to LayerNorm.
- Comparison to Batch Normalization: LayerNorm normalizes across the features of a single data point, whereas Batch Normalization normalizes across the batch dimension for each feature. LayerNorm is generally preferred in transformers due to its independence from batch size and better empirical performance.
2.3. Attention Mechanism Variations
- Computational Complexity: The standard self-attention has an
O(n^2)complexity with respect to sequence lengthn, which becomes prohibitive for long sequences. - Sliding Window Attention:
- Restricts each token to attend only to a fixed-size window of neighboring tokens.
- Reduces complexity to
O(n * w), wherewis the window size. - Often interleaved with global attention layers in modern models.
- Conceptually similar to the receptive field in computer vision convolutions.
- Shared Projection Matrices (Key/Value):
- Motivation: To reduce the number of parameters and memory usage, especially for the KV cache during decoding.
- Multi-Query Attention (MQA): All attention heads share a single set of projection matrices for keys and values.
- Group Query Attention (GQA): Projection matrices for keys and values are shared across groups of heads.
- Standard Multi-Head Attention: Each head has its own projection matrices for Q, K, and V.
- Rationale for Sharing K/V: Keys and values are accessed repeatedly during decoding (e.g., via KV cache), making parameter sharing here more impactful for memory efficiency. Queries, being more about "asking" questions, might benefit from diversity.
- Choice of Method: The selection of attention variation depends on factors like performance, latency requirements, model size, and input length. GQA is a common choice in recent models.
3. Evolution of Transformer Architectures
The lecture outlines the progression of transformer architectures from the original encoder-decoder to encoder-only and decoder-only models.
- Encoder-Decoder (Original Transformer, T5 family):
- T5 (Text-to-Text Transfer Transformer): Treats all NLP tasks as text-to-text problems. Uses a "span corruption" objective function during pre-training, where parts of the input are masked, and the decoder reconstructs them.
- mT5: Multilingual version of T5.
- ByT5: Byte-level tokenizer-free version, operating on bytes instead of subword tokens, leading to a much smaller vocabulary.
- Encoder-Only (BERT, DistilBERT, RoBERTa):
- Concept: Utilizes only the encoder stack, discarding the decoder. This architecture is well-suited for tasks requiring rich contextual representations, such as classification and sequence labeling, but not for generation.
- BERT (Bidirectional Encoder Representations from Transformers):
- Bidirectionality: Achieved by using the encoder's self-attention without masking, allowing each token to attend to all other tokens (both preceding and succeeding). This contrasts with GPT's unidirectional attention.
- Pre-training Objectives:
- Masked Language Model (MLM): Randomly masks tokens and trains the model to predict them based on context. Masking strategies include replacing with
[MASK](80%), keeping the original token (10%), or replacing with a random token (10%). - Next Sentence Prediction (NSP): Trains the model to predict whether two input sentences are consecutive in the original corpus. Uses
[CLS]token for classification and[SEP]tokens to separate sentences.
- Masked Language Model (MLM): Randomly masks tokens and trains the model to predict them based on context. Masking strategies include replacing with
- Input Representation: Combines token embeddings, positional embeddings, and segment embeddings (for NSP).
- Fine-tuning: The pre-trained encoder is used, and a task-specific layer (e.g., a linear classifier on the
[CLS]token output) is added and trained. - Pros: Excellent for classification tasks, leverages unlabeled data effectively, requires less data for fine-tuning.
- Cons: Not suitable for text generation, can be computationally expensive.
- DistilBERT: A distilled version of BERT, significantly smaller and faster by reducing the number of layers and using distillation to retain performance.
- RoBERTa: An optimized BERT variant that removes the NSP objective, uses dynamic masking, and trains on a larger dataset with more diverse data strategies.
- Decoder-Only (GPT family):
- Concept: Utilizes only the decoder stack, primarily for generative tasks. The masked self-attention in the decoder ensures causality (tokens only attend to preceding tokens).
- Evolution: Modern LLMs are largely decoder-only, as next-word prediction is a simple yet effective objective that scales well and aligns with chatbot applications. The compute budget is often better invested in the decoder.
4. Deep Dive into BERT
- Acronym Breakdown:
- Bidirectional: Attends to tokens both before and after the current token.
- Encoder: Uses the encoder part of the transformer architecture.
- Representations: Aims to learn rich, contextualized embeddings.
- Transformers: Based on the transformer architecture.
- Key Innovations:
- Bidirectional Context: Achieved by removing the causal mask in self-attention.
- Pre-training Objectives (MLM & NSP): Designed to learn general language understanding capabilities.
- Input Structure: Introduction of
[CLS]and[SEP]tokens, and segment embeddings. - WordPiece Tokenization: A subword tokenization method learned from data.
- Training Stages:
- Pre-training: On large unlabeled datasets using MLM and NSP.
- Fine-tuning: On specific downstream tasks with labeled data.
- MLM Details:
- Randomly masks tokens (80%
[MASK], 10% original, 10% random). - The model predicts the original masked tokens.
- Randomly masks tokens (80%
- NSP Details:
- Takes pairs of sentences.
- Predicts if sentence B follows sentence A (classification task).
- Uses
[CLS]token for classification and[SEP]to separate sentences.
- Input Embeddings: Sum of token embeddings, positional embeddings, and segment embeddings.
- FFN in BERT: A feed-forward network within each encoder layer, typically a two-layer MLP, used for further processing after attention.
- Classification Task Example (Sentiment Extraction):
- Input sentence is tokenized, prepended with
[CLS], and appended with[SEP]. [PAD]tokens are used for batching.- Embeddings are computed (token + position + segment).
- The sequence passes through the BERT encoder.
- The output embedding of the
[CLS]token is fed into a linear layer for classification. Other token embeddings are discarded for this specific task.
- Input sentence is tokenized, prepended with
- Limitations of BERT:
- Context Length: Original BERT had a limited context window (e.g., 512 tokens).
- Latency/Cost: Large parameter count (e.g., 110 million for BERT-base) leads to high latency.
- Pre-training Objectives: Questions about the necessity and effectiveness of both MLM and NSP.
- Addressing Limitations:
- Distillation (DistilBERT): Smaller, faster models trained to mimic the output distribution of larger models. Uses KL divergence as a loss.
- RoBERTa: Improves performance by removing NSP, using dynamic masking, and training on more data.
5. Conclusion and Future Outlook
The lecture highlights the continuous evolution of the transformer architecture, driven by the need for greater efficiency, scalability, and performance across various NLP tasks. Key takeaways include the importance of positional encoding, the shift towards pre-norm LayerNorm, the development of efficient attention mechanisms, and the specialization of architectures into encoder-only, decoder-only, and encoder-decoder models. Modern LLMs predominantly leverage decoder-only architectures for their scalability and effectiveness in next-word prediction.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Every Kind of Volcano | SciShow Kids
SciShow Kids

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth