Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 8: Parallelism

Stanford OnlineAbout 4 min readApr 29, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Parallelism Strategies: Data Parallelism (DDP), Tensor Parallelism (TP), Pipeline Parallelism (PP), Expert Parallelism (EP), Sequence/Context Parallelism.
  • Communication Primitives: All-Reduce, Reduce-Scatter, All-Gather.
  • Memory Management: ZeRO (Zero Redundancy Optimizer) stages 1, 2, and 3 (FSDP).
  • Hardware Topologies: Toroidal Mesh (TPU) vs. Fat Tree (GPU).
  • Performance Metrics: GPU Utilization, Communication-to-Computation ratio, "Bubbles" (idle time in pipelines).
  • Optimization Techniques: Activation Recomputation (Gradient Checkpointing), Flash Attention, Zero-Bubble Pipelining.

1. Hardware and Networking Foundations

The lecture emphasizes that the "unit of compute" is no longer a single GPU, but the entire data center.

  • TPUs vs. GPUs: TPUs utilize a toroidal mesh topology, allowing for efficient neighbor-to-neighbor communication, which scales well for dense models. GPUs typically use a fat tree topology, which is more flexible for "all-to-all" communication patterns, essential for modern Mixture of Experts (MoE) models.
  • Hardware Trade-offs: The Huawei Ascend 910 demonstrates a "brute force" approach—using slower individual chips but connecting a massive number (384+) via high-speed fiber optics, albeit at a significantly higher power cost compared to Nvidia systems.

2. Data Parallelism and ZeRO Stages

Data parallelism is the baseline, but it is memory-inefficient because every GPU stores a full copy of the model.

  • ZeRO Stage 1: Shards only the optimizer state across GPUs. Communication cost is equivalent to standard DDP (one All-Reduce).
  • ZeRO Stage 2: Shards both optimizer states and gradients.
  • ZeRO Stage 3 (FSDP): Shards parameters, gradients, and optimizer states. It uses an "on-demand" approach: gathering weights for a layer, computing, and immediately discarding them. While it requires extra communication, this is hidden by overlapping computation and communication.

3. Model Parallelism Strategies

When models exceed the memory capacity of a single node, model parallelism is required:

  • Pipeline Parallelism (PP): Splits the model by layers. It suffers from "bubbles" (idle time) where GPUs wait for activations. Zero-bubble pipelining mitigates this by separating the backward propagation of partial derivatives from the weight gradient computation.
  • Tensor Parallelism (TP): Splits individual matrix multiplications (matmuls) across GPUs. It is highly communication-intensive and is best suited for fast intra-node interconnects (e.g., NVLink).
  • Expert Parallelism (EP): Used for MoE models. It shards the Feed-Forward Network (FFN) layers. It is generally preferred over TP for MoE models because it is more efficient at routing sparse token activations.

4. Activation Memory and Sequence Parallelism

Activations often dwarf parameter memory in large models.

  • Activation Recomputation: A trade-off where memory is saved by re-calculating activations during the backward pass instead of storing them.
  • Sequence Parallelism: Splits activations along the sequence dimension. This is crucial for long-context models and is often used in conjunction with TP to reduce the memory footprint of layer norms and other non-sharded operations.

5. Practical Implementation Guidelines

The lecturer provides a "practitioner’s prescription" for configuring large-scale training:

  1. Maximize Data Parallelism: Use as many GPUs as possible for data sharding.
  2. Intra-node (Fast Links): Use Tensor Parallelism or Expert Parallelism (limit TP to 8 GPUs per node to stay within NVLink).
  3. Inter-node (Slow Links): Use Pipeline Parallelism to bridge multiple nodes.
  4. MoE Preference: If training an MoE, prioritize Expert Parallelism over Tensor Parallelism.
  5. Long Context: Use Context Parallelism (Ring Attention) for memory-efficient long-sequence processing.

6. Notable Quotes

  • "The new unit of compute is not the GPU, it's the entire data center."
  • "There is no one strictly dominant parallelization strategy. It's all a whole bunch of trade-offs that you somehow have to manage gracefully."
  • "If you're going to do some sort of parallelism, either EP or TP, you should be using EP over TP [for MoE models]."

Synthesis

Modern large-scale training is a complex orchestration of multiple parallelization strategies. The goal is to keep the compute units fully utilized by balancing memory constraints (via sharding/ZeRO) and communication overhead (via topology-aware placement). By combining these techniques—Data, Tensor, Pipeline, and Expert parallelism—engineers can scale models to hundreds of billions of parameters while maintaining high hardware efficiency. The "optimal" configuration is highly dependent on the model architecture (Dense vs. MoE) and the underlying network topology.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.