Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 3 - Tranformers & Large Language Models
By Unknown Author
Key Concepts
- Large Language Models (LLMs): Models trained on massive datasets with billions of parameters, capable of text-to-text generation.
- Transformer Architecture: The foundational architecture for modern LLMs, comprising encoder-decoder, encoder-only, and decoder-only variants.
- Decoder-Only Models: The dominant architecture for LLMs, focusing on generating text sequentially.
- Mixture of Experts (MoE): An architectural approach where different "experts" (sub-networks) handle different parts of the input, activated selectively.
- Sparse MoE: A variant of MoE that activates only a small subset (top K) of experts for efficiency.
- Routing Collapse: A training challenge in MoE models where the router consistently selects only a few experts, leading to underutilization.
- Next Token Prediction: The core task of LLMs, where the model predicts the most probable next token in a sequence.
- Greedy Decoding: A simple decoding strategy that always selects the token with the highest probability.
- Beam Search: A decoding strategy that explores multiple probable sequences (beams) to find a more globally optimal output.
- Sampling: A decoding strategy that randomly selects the next token based on its probability distribution.
- Top-K Sampling: A sampling method that restricts the selection to the K most probable tokens.
- Top-P (Nucleus) Sampling: A sampling method that selects tokens whose cumulative probability exceeds a threshold P.
- Temperature: A hyperparameter in softmax that controls the randomness of token selection; lower temperature leads to more deterministic output, higher temperature to more creative output.
- Prompting: The art of crafting input text to guide LLMs towards desired outputs.
- In-Context Learning: The ability of LLMs to learn from examples provided within the prompt without updating model weights.
- Zero-Shot Learning: Performing a task with no examples provided.
- Few-Shot Learning: Performing a task with a few examples provided.
- Chain-of-Thought (CoT) Prompting: Encouraging LLMs to generate intermediate reasoning steps before providing the final answer.
- Self-Consistency: A technique that involves sampling multiple responses and selecting the most frequent answer through majority voting.
- KV Caching: A technique to optimize inference by storing and reusing key and value matrices from previous computations.
- Group Query Attention (GQA): An attention mechanism that groups keys and values to reduce computational and memory overhead.
- Page Attention: A memory management technique for KV caching that reduces fragmentation by using fixed-size blocks.
- Multi-Latent Attention: An attention mechanism that factorizes projection matrices to create more compact representations for keys and values.
- Speculative Decoding: A technique that uses a smaller, faster "draft" model to generate candidate tokens, which are then verified by the larger LLM.
- Multi-Token Prediction: A training and inference strategy where the model predicts multiple tokens simultaneously.
Lecture 3: Large Language Models and Generation Strategies
This lecture introduces Large Language Models (LLMs) and delves into their architecture, generation mechanisms, and prompting strategies.
Recap of Transformer Architectures
Last week's lecture reviewed the three main categories of transformer-based models:
- Encoder-Decoder Models: Utilize both the encoder and decoder of the transformer. Typically used for text-in, text-out tasks (e.g., T5).
- Encoder-Only Models: Employ only the encoder of the transformer. Known for producing meaningful embeddings for tasks like classification and sentiment analysis (e.g., BERT). The embedding of the
CLStoken is often used for sentence-level representation. - Decoder-Only Models: Utilize only the decoder of the transformer, with cross-attention removed. These are text-in, text-out models and are the prevalent architecture for modern LLMs (e.g., GPT).
Introduction to Large Language Models (LLMs)
LLMs are language models that are "large" in three key aspects:
- Model Size: Typically have hundreds of billions of parameters, with a minimum threshold of around one billion parameters.
- Training Data: Trained on vast amounts of data, quantified by hundreds of billions or even trillions of tokens.
- Compute Requirements: Require significant computational resources, often involving numerous GPUs.
It's important to note that the definition of LLMs is relatively recent. Models like BERT, while powerful, are not considered LLMs under the current definition because they are encoder-only and do not produce text. Modern LLMs are predominantly decoder-only and perform text-to-text generation.
Mixture of Experts (MoE) Architecture
To address the computational cost of activating all parameters in large models for every inference, the Mixture of Experts (MoE) architecture is introduced.
- Concept: Instead of engaging all parameters, a "gate" or "router" network selects a subset of specialized "experts" (sub-networks) to process the input.
- Formula: The output
yis a weighted sum of expert outputs:y = sum(gate_output * expert_output). - Types of MoE:
- Dense MoE: All expert outputs are considered, with weights determining their contribution.
- Sparse MoE: Only the top-K experts are selected and activated, significantly reducing computation. K is a hyperparameter (e.g., K=1 or K=2).
- Expert Placement: In LLMs, MoE is typically implemented within the Feed-Forward Network (FFN) layers, as these layers are parameter-heavy and contribute significantly to computation.
- Training Challenges:
- Routing Collapse: The router may consistently select only a few experts, leading to underutilization of others.
- Mitigation: An auxiliary loss term is added to the training objective to encourage more uniform expert utilization. This loss incentivizes the fraction of tokens routed to each expert (
f_i) and the average routing probability for each expert (p_i) to converge towards uniform distributions. - Noisy Gating: Adding noise to the gate outputs can help activate different experts by chance.
- Benefits: MoE allows for scaling model capacity (more parameters) without a proportional increase in compute at inference time. This can lead to more sample-efficient training.
- Implementation: Experts are typically FFNs, and routing is done at the token level, meaning each token can be routed to a different expert within a layer. The router (gate) is layer-specific and trainable.
Response Generation Strategies
Once an LLM is trained, generating a response involves predicting the next token. Several strategies exist:
- Greedy Decoding: Always selects the token with the highest probability.
- Limitations: Leads to deterministic and repetitive outputs, and can be locally optimal but not globally optimal for sequence generation.
- Beam Search: Maintains a fixed number (beam width, K) of the most probable sequences at each step.
- Goal: To find a more globally optimal sequence than greedy decoding.
- Limitations: Can still lack diversity and creativity. Prioritizes shorter sequences due to the multiplicative nature of probabilities.
- Sampling: Randomly selects the next token based on its probability distribution.
- Top-K Sampling: Restricts sampling to the K most probable tokens.
- Top-P (Nucleus) Sampling: Selects tokens whose cumulative probability exceeds a threshold P.
- Temperature (T): A hyperparameter that controls the shape of the probability distribution.
- Low Temperature: Creates a "spiky" distribution, favoring high-probability tokens (more deterministic, less creative).
- High Temperature: Creates a more uniform distribution, increasing the likelihood of sampling less probable tokens (more creative, less deterministic).
- T=0: Theoretically leads to deterministic output (greedy decoding), but practical implementations may still exhibit non-determinism due to hardware optimizations.
Obtaining Probabilities: The LLM's final layers, typically a linear layer followed by a softmax function, project the encoded representation into a probability distribution over the vocabulary.
Guided Decoding
For generating output in specific formats (e.g., JSON), Guided Decoding can be used. This technique filters out invalid next tokens during the generation process, ensuring adherence to structural constraints. This can be implemented using finite state machines or context grammars.
Prompting Strategies and In-Context Learning
Effective prompting is crucial for eliciting desired responses from LLMs.
- Prompt Structure: Prompts typically include:
- Context: Setting the scene or providing background information.
- Instructions: Explicit commands for the LLM to perform a task.
- Input: The specific data the LLM should process.
- Constraints: Rules or limitations for the output.
- Context Length: The maximum number of tokens an LLM can process at once. Modern LLMs can handle tens of thousands to millions of tokens.
- Context Rot: A phenomenon where LLM performance degrades with increasing context length, particularly in information retrieval tasks, due to distractors and model limitations.
- In-Context Learning: LLMs can learn from examples provided within the prompt without weight updates.
- Zero-Shot: No examples are provided; the LLM relies solely on its pre-training.
- Few-Shot: A few input-output examples are provided to guide the LLM.
- Chain-of-Thought (CoT) Prompting: Encouraging the LLM to generate intermediate reasoning steps before the final answer. This improves performance, especially for complex tasks, and aids in debugging by making the model's reasoning transparent.
- Self-Consistency: Enhancing CoT by sampling multiple responses, extracting the final answer from each, and using majority voting to determine the most robust answer. This can be parallelized for efficiency.
Inference Efficiency Techniques
Optimizing LLM inference is critical due to their size and computational demands.
-
Exact Methods (Optimizing Computations):
- KV Caching: Storing and reusing key and value matrices from previous token computations to avoid redundant calculations.
- Group Query Attention (GQA): Grouping keys and values across attention heads to reduce memory bandwidth and computation.
- Page Attention: A memory management system that uses fixed-size blocks to store KV cache, reducing fragmentation and improving memory utilization.
- Multi-Latent Attention: Factorizing projection matrices for keys and values into lower-dimensional spaces, and sharing these matrices across heads, leading to more compact representations and potential regularization benefits.
-
Approximate Methods (Using Approximations):
- Speculative Decoding: Using a smaller, faster "draft" model to generate candidate tokens, which are then verified by the larger LLM. This allows for faster generation by processing multiple tokens in parallel. The acceptance/rejection mechanism ensures the output distribution matches the target model.
- Multi-Token Prediction: Modifying the model architecture and training objective to predict multiple tokens simultaneously, potentially embedding the draft model within the main model.
The lecture concludes by emphasizing that these techniques offer trade-offs between efficiency, accuracy, and complexity, and encourages further exploration of the referenced research papers.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth

Nvidia Crushes Earnings again — What Jensen Huang sees next for AI
CGTN America

Samsung union suspends strike after reaching tentative pay deal • FRANCE 24 English
FRANCE 24 English

OH SH*T! The Banks are Dumping AI Loans!
Steven Van Metre