Stanford CS229 I Machine Learning I Building Large Language Models (LLMs)

Stanford OnlineAbout 9 min readFeb 1, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • LLMs (Large Language Models): Chatbots like ChatGPT, Claude, Gemini, and Llama.
  • Pretraining: Training a language model to model all of the internet.
  • Post-training: Taking large language models and making them AI assistants.
  • Autoregressive Language Models: Models that predict the next word based on the preceding words in a sequence.
  • Tokenization: The process of breaking down text into smaller units (tokens) for processing by language models.
  • Byte Pair Encoding (BPE): A common tokenization algorithm that merges frequent pairs of characters or tokens.
  • Perplexity: A metric used to evaluate language models, representing the number of tokens the model is "hesitating" between when predicting the next word.
  • HELM & Hugging Face Open Leaderboard: Common benchmarks for evaluating LLMs on various NLP tasks.
  • MMLU (Massive Multitask Language Understanding): A benchmark consisting of question-answering tasks across various domains.
  • Train-Test Contamination: The issue of test data being present in the training data, leading to inflated performance metrics.
  • Scaling Laws: The observation that the performance of language models improves predictably with increased data and model size.
  • Chinchilla Scaling Laws: A set of scaling laws that determine the optimal allocation of training resources (data and model size) for language models.
  • SFT (Supervised Fine-Tuning): Fine-tuning a pre-trained language model on a dataset of desired answers collected from humans.
  • RLHF (Reinforcement Learning from Human Feedback): A technique for aligning language models with human preferences using reinforcement learning.
  • Reward Model: A classifier trained to predict human preferences between different outputs from a language model.
  • PPO (Proximal Policy Optimization): A reinforcement learning algorithm used in RLHF to fine-tune language models based on human feedback.
  • DPO (Direct Preference Optimization): A simplified alternative to PPO that directly optimizes the language model based on human preferences.
  • Model FLOP Utilization: A metric that measures the efficiency of GPU usage during model training.
  • Automatic Mixed Precision (AMP): A technique that uses both 32-bit and 16-bit floating-point numbers to speed up training and reduce memory consumption.
  • Operator Fusion: A system optimization technique that combines multiple operations into a single kernel to reduce communication overhead.

1. Introduction to Building LLMs

  • LLMs (Large Language Models) are the foundation of chatbots like ChatGPT, Claude, Gemini, and Llama.
  • The lecture provides an overview of the key components involved in training LLMs.
  • Five key components matter when training LLMs: architecture, training loss and algorithm, data, evaluation, and systems.
  • The lecture focuses on data, evaluation, and systems, as these are often overlooked in academic research but are crucial in practice.
  • The lecture covers pretraining (modeling all of the internet) and post-training (making AI assistants).

2. Pretraining LLMs: Language Modeling

  • Task: Language modeling, which involves training a model to predict the probability distribution over sequences of tokens or words.
  • Language Model Definition: A model that assigns a probability to a given sentence, reflecting its likelihood of being uttered by a human or found online.
    • Example: "The mouse ate the cheese" should have a higher probability than "The cheese ate the mouse."
  • Generative Models: Models that can generate new data (e.g., sentences) by sampling from their learned probability distribution.
  • Autoregressive Language Models: A type of language model that decomposes the probability of a sequence into the product of conditional probabilities of each word given the preceding words.
    • Formula: P(x1, x2, ..., xL) = P(x1) * P(x2|x1) * P(x3|x1, x2) * ... * P(xL|x1, ..., xL-1)
    • Sampling from autoregressive models involves a for loop, generating one word at a time.
  • Autoregressive Language Model Task: Predicting the next word in a sequence.
    • Process: Tokenize the input sequence, pass it through a model, obtain a probability distribution over the next token, sample from the distribution, and detokenize the output.
    • Training only requires predicting the most likely token and comparing it to the actual token.
  • Autoregressive Neural Language Models:
    • Embed each token into a vector representation.
    • Pass the embeddings through a neural network (e.g., a transformer).
    • Apply a linear layer to map the output to the size of the vocabulary.
    • Use a softmax function to obtain a probability distribution over the next words.
    • Use cross-entropy loss to train the model to predict the correct next token.
    • Minimizing the cross-entropy loss is equivalent to maximizing the log-likelihood of the text.

3. Tokenization

  • Importance: Tokenizers are crucial for processing text in LLMs.
  • Why Tokenizers are Needed:
    • More general than words: Handles typos and languages without spaces between words (e.g., Thai).
    • Avoids super-long sequences: Character-by-character tokenization would result in excessively long sequences, increasing computational complexity.
  • Tokenizers: Give common subsequences a certain token. The average token is around 3-4 letters.
  • Byte Pair Encoding (BPE): A common tokenization algorithm.
    • Process: Start with a large corpus of text, assign each character a unique token, and iteratively merge the most frequent pairs of tokens.
  • Pre-tokenizers: Handle spaces and punctuation separately for computational efficiency.
  • Token Uniqueness: Each token has a unique ID.
  • Drawbacks of Tokenizers:
    • Numbers are not tokenized, making it difficult for models to generalize in math.
    • Special tokenization is needed for code.
  • Future Trends: Moving away from tokenizers towards character-by-character or byte-by-byte tokenization as architectures improve.

4. Evaluation: Perplexity and Benchmarks

  • Perplexity: A metric used to evaluate language models.
    • Formula: Perplexity = 2^(average per-token loss).
    • Range: Between 1 and the size of the vocabulary.
    • Interpretation: The number of tokens the model is "hesitating" between.
    • Limitations: Depends on the tokenizer and the evaluation data.
  • NLP Benchmarks: Evaluating LLMs on classical NLP tasks.
    • Examples: HELM (Stanford), Hugging Face Open Leaderboard.
    • Tasks: Question answering, where the model's likelihood of generating the correct answer is compared to other answers.
  • MMLU (Massive Multitask Language Understanding): A common academic benchmark.
    • Consists of question-answering tasks across various domains (e.g., college medicine, physics, astronomy).
    • Evaluation involves assessing the model's likelihood of generating the correct answer from a set of options.
  • Evaluation Challenges:
    • Inconsistencies in evaluation methods across different organizations.
    • Train-test contamination: Test data being present in the training data.

5. Data: Pretraining Data Collection and Processing

  • Data Source: "All of the internet" (or "clean internet").
  • Process:
    1. Web Crawling: Downloading web pages using web crawlers (e.g., Common Crawl).
      • Common Crawl contains around 250 billion pages (1 petabyte of data).
    2. Text Extraction: Extracting text from HTML, handling math, and removing boilerplates.
    3. Filtering Undesirable Content: Removing not-safe-for-work, harmful content, and PII (Personally Identifiable Information).
    4. De-duplication: Removing duplicate content (e.g., headers, footers, repeated paragraphs).
    5. Heuristic Filtering: Removing low-quality documents based on rules (e.g., outlier tokens, word length).
    6. Model-Based Filtering: Training a classifier to identify high-quality documents based on references from Wikipedia.
    7. Domain Classification: Classifying data into different domains (e.g., entertainment, books, code) and up/down-weighting them.
    8. High-Quality Data Training: Training on very high-quality data (e.g., Wikipedia) at the end of training with a decreased learning rate.
  • Data is Key: Collecting well data is a huge part of practical, large language model.
  • Data Processing Challenges:
    • Efficient processing of large datasets.
    • Balancing different domains.
    • Exploring synthetic data generation.
    • Utilizing multimodal data.
  • Data Secrecy: Companies often keep their data collection methods secret due to competitive dynamics and copyright liability issues.
  • Dataset Sizes:
    • Academic benchmarks: Started from around 150 billion tokens (800 GB) to 15 trillion tokens.
    • Llama 2: Trained on 2 trillion tokens.
    • Llama 3: Trained on 15 trillion tokens.

6. Scaling Laws

  • Observation: The more data and the larger the models, the better the performance.
  • Scaling Laws: Predict how much better the performance will be if the amount of data and the size of the model are increased.
  • Linear Relationship (Log Scale): The relationship between compute, data, parameters, and test loss is linear when plotted on a log scale.
  • Practical Implications:
    • Predicting future performance based on increased compute.
    • Optimizing model training by finding a scaling recipe.
    • Tuning hyperparameters on smaller models and extrapolating to larger models.
  • Scaling Rate and Intercept: Important factors in scaling laws.
  • Optimal Allocation of Training Resources:
    • Chinchilla paper: Determined the optimal number of parameters and tokens for a given amount of compute.
    • Optimal ratio: 20 tokens per parameter for training resources.
    • Considering inference costs: 150 tokens per parameter.
  • The Bitter Lesson: The importance of architectures that can leverage computation, emphasizing systems and data over minor architectural differences.

7. Back of the Envelope Computation (Llama 3 Example)

  • Llama 3 400B:
    • Trained on 15.6 trillion tokens.
    • 405 billion parameters.
  • Flops Calculation: 6 * (number of parameters) * (number of tokens) = 3.8e25 flops.
  • Hardware: 16,000 H100s.
  • Training Time: Approximately 70 days (26 million GPU hours).
  • Cost:
    • GPU rental: $52 million (assuming $2 per hour per H100).
    • Salaries (50 employees): $25 million.
    • Total: Approximately $75 million.
  • Carbon Emission: Approximately 4000 tons of CO2 equivalent.

8. Post-Training: Alignment and RLHF

  • Task: Making AI assistants that follow instructions and align with human values.
  • Motivation: Language modeling alone is insufficient for AI assistants.
  • Process: Fine-tuning a pre-trained language model on data that reflects desired behavior.
  • SFT (Supervised Fine-Tuning): Fine-tuning the LLM on desired answers collected from humans.
    • Collecting Data: Asking humans to provide questions and answers.
    • Scaling Data Collection: Using LLMs to generate more question-answer pairs.
    • Quantity of Data: SFT doesn't require much data.
  • RLHF (Reinforcement Learning from Human Feedback): Aligning language models with human preferences using reinforcement learning.
    • Maximizing Human Preference: Instead of cloning human behavior, RLHF maximizes human preference.
    • Process: Generate two answers for a given instruction, ask labelers to select the preferred one, and fine-tune the model to generate more of the preferred output.
    • Reward Model: Training a classifier to predict human preferences between different outputs.
    • PPO (Proximal Policy Optimization): A reinforcement learning algorithm used to fine-tune the language model based on the reward model.
    • DPO (Direct Preference Optimization): A simplified alternative to PPO that directly optimizes the language model based on human preferences.
  • Data Collection for RLHF:
    • Humans: Slow, expensive, and prone to biases.
    • LLMs: Cheaper and can achieve higher agreement with the mode of humans.
  • Evaluation of Post-Training:
    • Challenges: Cannot use validation loss or perplexity.
    • ChatBotArena: A benchmark where random users interact with two chatbots and rate which one is better.
    • AlpacaEval: A benchmark that uses LLMs to evaluate model outputs.

9. Systems: Optimizing Compute

  • Compute Bottleneck: Compute is a major bottleneck in training LLMs.
  • GPU Optimization:
    • GPUs are optimized for throughput and fast matrix multiplication.
    • Compute has been improving faster than memory and communication.
    • Model FLOP Utilization: A metric that measures the efficiency of GPU usage.
  • Low Precision: Using 16-bit floating-point numbers instead of 32-bit to speed up training and reduce memory consumption.
  • Automatic Mixed Precision (AMP): Using both 32-bit and 16-bit floating-point numbers.
  • Operator Fusion: Combining multiple operations into a single kernel to reduce communication overhead.

10. Synthesis/Conclusion

The lecture provides a comprehensive overview of the key components involved in building large language models, emphasizing the importance of data, evaluation, and systems. It covers pretraining techniques, tokenization methods, evaluation metrics, data collection and processing strategies, scaling laws, post-training alignment techniques (SFT and RLHF), and system optimization strategies. The lecture highlights the practical challenges and trade-offs involved in training LLMs and emphasizes the importance of optimizing compute resources and data quality.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.