Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 1 - Transformer
By Unknown Author
Key Concepts
- Natural Language Processing (NLP): The field of computing focused on manipulating and understanding text.
- NLP Tasks:
- Classification: Predicting a single label from text (e.g., sentiment analysis, intent detection).
- Multi-classification: Predicting multiple labels from text (e.g., Named Entity Recognition).
- Generation: Producing text as output based on text input (e.g., machine translation, question answering, summarization).
- Tokenization: The process of breaking down text into smaller units called tokens.
- Word-level tokenization: Separating text by words.
- Subword tokenization: Breaking words into smaller meaningful units (roots, prefixes, suffixes) to handle variations and reduce Out-Of-Vocabulary (OOV) issues.
- Character-level tokenization: Breaking text into individual characters, robust to misspellings but leading to very long sequences.
- Out-Of-Vocabulary (OOV): Tokens encountered during inference that were not seen during training.
- Token Representation/Embeddings: Numerical representations of tokens that models can understand.
- One-Hot Encoding (OHE): A naive method where each token is represented by a vector with a single '1' and the rest '0's. Lacks similarity information.
- Word2vec: A pioneering technique for learning word embeddings using proxy tasks (Continuous Bag-of-Words and Skip-gram) to capture semantic relationships.
- Cosine Similarity: A measure of similarity between two non-zero vectors, often used to compare token embeddings.
- Recurrent Neural Networks (RNNs): Models that process sequences by maintaining a hidden state that captures information from previous tokens.
- Long Short-Term Memory (LSTM): An extension of RNNs designed to better handle long-range dependencies and mitigate the vanishing gradient problem.
- Vanishing Gradient: A problem in training deep neural networks where gradients become very small, hindering learning in earlier layers.
- Attention Mechanism: A mechanism that allows models to focus on specific parts of the input sequence when processing or generating output, addressing long-range dependency issues.
- Transformer Architecture: A neural network architecture that relies heavily on the attention mechanism, particularly self-attention, to process sequences in parallel.
- Self-Attention: Allows tokens within the same sequence to attend to each other, creating context-aware representations.
- Encoder-Decoder Structure: The Transformer typically consists of an encoder to process input and a decoder to generate output.
- Query (Q), Key (K), Value (V): Components used in the attention mechanism to calculate attention weights and derive context-aware representations.
- Multi-Head Attention: Performing the attention mechanism multiple times in parallel with different learned projections to capture diverse relationships.
- Positional Encoding: Information added to token embeddings to indicate their position in the sequence, as self-attention itself is order-agnostic.
- Masked Self-Attention: A variant used in the decoder that prevents attention to future tokens in the output sequence.
- Cross-Attention: Allows the decoder to attend to the output of the encoder.
- Feedforward Network (FFN): A standard neural network layer applied after attention layers in the Transformer.
- Label Smoothing: A regularization technique that slightly modifies target labels to prevent overconfidence in predictions.
Course Introduction and Logistics
The class, CME 295: Transformers and Large Language Models, is taught by twin brothers Afshine and Shervine, who have industry experience at Uber, Google, and Netflix, specializing in NLP and LLMs. The course aims to cover the underlying mechanisms of LLMs, focusing on the Transformer architecture, and how these models are trained and applied.
Target Audience: Individuals interested in NLP and LLMs, aspiring research or ML scientists, developers building LLM-powered projects, or those seeking to understand AI/GenAI applications in their domain.
Prerequisites: Foundational knowledge in Machine Learning (model training, neural networks) and basic Linear Algebra (matrix multiplication).
Schedule and Format:
- Held every Friday from 3:30 PM to 5:20 PM.
- Two units, available for letter or credit/non-credit grading.
- Lectures are recorded and made available online.
Assessment:
- Two exams: Midterm (October 24) and Final (week of December 8, date TBD).
- Exams focus on concepts, not coding.
- Exam weighting: 50% Midterm, 50% Final. No homework.
- The final exam will likely cover the second half of the course topics.
Resources:
- Slides and recordings posted on the course website.
- Syllabus available on the website.
- Textbook: "Super Study Guide-- Transformer LLMs."
- "VIP cheat sheet" available on GitHub, translated into multiple languages.
Communication:
- Announcements on Canvas.
- Questions can be posted on the Ed tab on Canvas.
- Contact via mailing list or direct message.
- Instructors will repeat questions asked during lectures for clarity in recordings.
Waitlist: For waitlisted students, it's recommended to talk to the instructors; the current waitlist size (six) suggests a high probability of enrollment.
Natural Language Processing (NLP) Overview
NLP is the field of computing that deals with manipulating and processing text. Tasks are broadly categorized into three buckets:
1. Classification
- Description: Input text is used to predict a single output.
- Examples:
- Sentiment Analysis: Determining if a movie review is positive, negative, or neutral.
- Intent Detection: Identifying the user's goal (e.g., "create an alarm").
- Language Detection: Identifying the language of a text.
- Topic Modeling: Assigning topics to text.
- Evaluation Metrics:
- Accuracy: Percentage of correct predictions.
- Precision: Of all positive predictions, what percentage were correct?
- Recall: Of all true positive instances, what percentage were correctly predicted?
- F1 Score: Harmonic mean of precision and recall, useful for imbalanced datasets.
- Imbalanced Datasets: Accuracy can be misleading when class distribution is uneven (e.g., 99% positive, 1% negative). Precision and recall are crucial in such cases.
2. Multi-classification
- Description: Input text is used to predict multiple outputs simultaneously.
- Examples:
- Named Entity Recognition (NER): Identifying and categorizing specific words (e.g., locations, times, organizations).
- Part-of-Speech Tagging: Identifying nouns, verbs, adjectives, etc.
- Parsing: Analyzing the grammatical structure of sentences (dependency or constituency parsing).
- Evaluation: Metrics are typically applied at the token or entity-type level.
3. Generation
- Description: Text is generated as output based on text input, with variable output length. This is currently the most popular category.
- Examples:
- Machine Translation: Translating text from one language to another (e.g., English to German).
- Question Answering: Providing answers to questions (e.g., ChatGPT, Gemini).
- Summarization: Condensing longer texts into shorter summaries.
- Text Generation: Creating code, poems, or other forms of text.
- Evaluation Metrics (for Machine Translation):
- BLEU (Bilingual Evaluation Understudy): Compares generated translation to reference translations.
- ROUGE: A suite of metrics for evaluating summaries and translations.
- Perplexity: Measures how surprised a model is by its own output (lower is better).
- Challenges: Evaluating generation tasks is complex due to the subjectivity of "correct" output. Reference-based metrics require costly labeled data. The field is moving towards reference-free metrics.
Evolution of NLP Models
While LLMs gained prominence in 2022, the field has a longer history:
- 1980s: Early models were conceptualized.
- 1990s: LSTMs emerged.
- Limitations: Lack of internet and computational power hindered early model development.
- 2010s:
- Word2vec (around 2013): Pioneered meaningful word embeddings.
- Transformers (2017): Introduced the foundational architecture for modern LLMs.
- 2020s: Scaling up Transformer models with increased compute and data led to the development of current LLMs.
Text Representation: Tokenization and Embeddings
Models understand numbers, not text, so text must be converted into a numerical format.
Tokenization Methods:
- Word-level: Simple, but struggles with word variations (e.g., "bear" vs. "bears") and leads to large vocabularies and OOV issues.
- Subword: Leverages word roots to handle variations and reduce OOV risk. However, it can lead to longer sequences, increasing computational complexity.
- Character-level: Robust to misspellings but results in very long sequences, significantly slowing down computation.
Token Representation (Embeddings):
- One-Hot Encoding (OHE): Assigns a unique vector to each token. Problematic because it treats all tokens as equally dissimilar (orthogonal vectors) and doesn't capture semantic relationships.
- Learning Embeddings: The goal is to learn embeddings where similar tokens have similar representations.
- Word2vec:
- Proxy Tasks: Uses tasks like predicting a target word from its context (Continuous Bag-of-Words) or predicting context words from a target word (Skip-gram).
- Objective: The primary goal is to learn meaningful word representations, not necessarily to excel at the proxy task itself.
- Process: A simple neural network takes a one-hot encoded word as input, passes it through a hidden layer (embedding layer), and predicts output probabilities. The weights of the hidden layer become the word embeddings.
- Vocabulary Size (V): The number of unique tokens.
- Hidden Layer Dimension (d): The dimensionality of the learned embeddings, typically much smaller than V (e.g., 100-768).
- Training: Uses backpropagation to minimize a loss function (e.g., cross-entropy) by comparing predictions to true labels, updating weights to improve predictions.
- Stopping Criteria: Training stops when the loss converges or based on downstream task performance.
- Generation Stop: Special "end of sequence" (EOS) tokens signal the end of generated text.
- Contextual Embeddings: Word2vec embeddings are static and do not account for context. For example, "bank" in "riverbank" and "robbing a bank" would have the same embedding.
- Word2vec:
Sequential Models: RNNs and LSTMs
To address the context-agnostic nature of word embeddings and incorporate word order, sequential models were developed.
Recurrent Neural Networks (RNNs):
- Mechanism: Process tokens one by one, maintaining a "hidden state" that summarizes the sequence processed so far.
- Input: Current token's embedding and the previous hidden state.
- Output: A new hidden state and potentially an output prediction.
- Pros: Captures word order and sequential information.
- Cons:
- Long-Range Dependencies: Difficulty remembering information from distant past tokens due to the vanishing gradient problem.
- Vanishing Gradient: Gradients become extremely small during backpropagation through time, making it hard to update weights for early tokens.
- Slow Computation: Sequential processing is inherently slow, especially for long sequences.
Long Short-Term Memory (LSTM):
- Improvement over RNNs: Introduces a "cell state" and gating mechanisms (input, forget, output gates) to better control information flow and retain important information over longer sequences.
- Goal: To mitigate the vanishing gradient problem and improve memory of past information.
- Still Limited: While better, LSTMs can still struggle with very long sequences and are computationally intensive.
The Rise of Attention and Transformers
The limitations of RNNs and LSTMs, particularly their sequential nature and difficulty with long-range dependencies, led to the development of attention mechanisms and the Transformer architecture.
Attention Mechanism (Introduced 2014):
- Concept: Allows a model to dynamically focus on specific parts of the input sequence when producing an output.
- Benefit: Directly addresses long-range dependency issues by creating direct links between relevant input and output elements.
Transformer Architecture (Introduced 2017 in "Attention is All You Need"):
- Core Idea: Replaces recurrence entirely with self-attention, enabling parallel processing of the entire input sequence.
- Self-Attention: Each token in a sequence can attend to all other tokens in the same sequence to compute a context-aware representation. This allows for unique representations of words like "bank" based on their context.
- Encoder-Decoder Structure:
- Encoder: Processes the input sequence (e.g., source language in translation) using self-attention layers to create rich, context-aware embeddings.
- Decoder: Generates the output sequence (e.g., target language) using masked self-attention (attending only to previously generated tokens) and cross-attention (attending to the encoder's output).
- Key Components:
- Input Embeddings: Convert tokens into dense vectors.
- Positional Encoding: Added to embeddings to provide information about token position, as self-attention is order-agnostic.
- Multi-Head Attention: Performs self-attention multiple times in parallel with different learned projections (Query, Key, Value matrices) to capture diverse relationships.
- Query (Q), Key (K), Value (V): Learned projections of the input embeddings. Q is compared to K to determine attention weights, which are then applied to V.
- Scaling: Division by the square root of the key dimension ($d_k$) normalizes dot products in attention.
- Concatenation and Projection: Outputs from multiple heads are concatenated and projected back to the original embedding dimension.
- Feedforward Network (FFN): A standard neural network layer applied after attention, often with a larger hidden dimension than input/output to increase model capacity.
- Layer Normalization and Residual Connections: Used within the encoder and decoder blocks to stabilize training and improve gradient flow.
- Translation Example Walkthrough:
- Tokenization: Input sentence ("a cute teddy bear is reading") is tokenized, including BOS (Beginning Of Sentence) and EOS (End Of Sentence) tokens.
- Input Embeddings + Positional Encoding: Tokens are converted to embeddings, and positional information is added.
- Encoder:
- Self-Attention: Tokens attend to each other to create context-aware representations. Multi-head attention is used.
- FFN: Further processing of the attention outputs.
- This process is repeated for multiple encoder layers.
- Decoder:
- Start Token (BOS): The decoder begins with the BOS token.
- Masked Self-Attention: The decoder attends to itself, but only to previously generated tokens (causal attention).
- Cross-Attention: The decoder attends to the encoder's output (keys and values) using its own current representation as the query. This links the input and output sequences.
- FFN: Further processing.
- This process is repeated for multiple decoder layers.
- Output Layer: A linear projection followed by a softmax function predicts the probability distribution over the vocabulary for the next token.
- Generation Loop: The predicted token is fed back into the decoder for the next step, continuing until the EOS token is generated.
- Label Smoothing: A technique where the target label is not a hard 1 for the correct word and 0s for others, but rather a slightly softened distribution (e.g., 1-$\epsilon$ for the correct word, $\epsilon/(V-1)$ for others). This encourages the model to be less overconfident and can improve metrics like BLEU.
The Transformer architecture, with its reliance on self-attention, has become the foundation for most modern Large Language Models.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Every Kind of Volcano | SciShow Kids
SciShow Kids

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth