Stanford CS336 Language Modeling from Scratch | Spring 2025 | Mixture of experts

Unknown AuthorAbout 8 min readApr 25, 2025Watch original
THE SUMMARYAI-generated

Mixture of Experts (MoE) - Deep Dive

Key Concepts:

  • Mixture of Experts (MoE): A sparsely activated architecture with multiple subcomponents ("experts") typically applied to the feed-forward network (FFN) layers of a transformer.
  • Experts: Multiple copies or split versions of the FFN layer.
  • Router: A mechanism that selects a subset of experts to process each token.
  • Sparsity: The characteristic of MoE models where only a fraction of the parameters (experts) are activated during each forward pass.
  • Token Choice: A routing strategy where each token selects its top-k preferred experts.
  • Expert Choice: A routing strategy where each expert selects its top-k preferred tokens.
  • Expert Parallelism: A parallelization strategy where each expert is placed on a separate device.
  • Fine-grained Experts: Smaller experts, achieved by reducing the hidden dimension of the FFN.
  • Shared Experts: Experts that are used by all tokens, regardless of the routing decision.
  • Auxiliary Loss: Additional loss terms used to balance the load across experts and devices.
  • Upcycling: A technique to convert a dense model into an MoE model by copying the FFN layers and adding a router.
  • MLHA (Multi-Head Latent Attention): An optimization technique used in DeepSeek V3 to compress the KV cache.
  • MTP (Multi-Token Prediction): A training technique used in DeepSeek V3 to predict multiple tokens in parallel.

1. Introduction to Mixture of Experts

  • MoE architectures are now prevalent in high-performance language models like Grok, DeepSeek, and Llama 4.
  • MoE models offer advantages over dense models in terms of performance for a given compute budget (flops).
  • The term "mixture of experts" is misleading; it doesn't imply specialized experts for different domains. Instead, it refers to a sparsely activated architecture.
  • The core idea is to replace the single FFN layer in a standard transformer with a router and multiple smaller or copied FFN layers (experts).

2. How Mixture of Experts Works

  • In a standard transformer, the feed-forward component is a single block. In an MoE model, this block is replaced with a selector layer (router) and multiple copies of the FFN.
  • The router selects a smaller number of experts for each token in each forward pass.
  • If only one expert is activated and its size is the same as the dense FFN, the flops are equivalent to a dense model, but the MoE model has more parameters.
  • The advantage is more parameters for memorizing facts without increasing computational cost.

3. Performance Benefits of Mixture of Experts

  • Multiple papers demonstrate that MoE models outperform dense models at the same flop count.
  • A paper by Fedus et al. (2022) shows that as the number of experts increases, the training loss decreases for a fixed flop count.
  • AI2's Olo paper confirms this, showing that MoE models have faster training loss decay compared to dense models.
  • DeepSeek V2 paper presents results showing that MoE models achieve better MMLU performance with fewer activated parameters.

4. Systems Benefits of Mixture of Experts

  • MoE allows for expert parallelism, where each expert can be placed on a different device.
  • The router directs tokens to the appropriate device for computation.
  • This is a natural way to shard large models across multiple devices.
  • Chinese groups like Quen and DeepSeek have been pioneers in open-source MoE implementations.

5. Challenges and Complexities of Mixture of Experts

  • Infrastructure is complex, and the biggest advantages are realized with multi-node training.
  • Routing decisions are non-differentiable, making optimization challenging.
  • Training objectives require careful engineering to ensure stability.

6. Routing Mechanisms

  • Token Choice: Each token selects its top-k preferred experts. This is the most common approach.
  • Expert Choice: Each expert selects its top-k preferred tokens. This ensures balanced utilization of experts.
  • Global Assignment: Solves a complex optimization problem to balance the mapping between experts and tokens.
  • Token choice with top-k routing is the dominant approach in practice.

7. Top-K Routing in Detail

  • The router computes a score for each token-expert pair, similar to an attention operation.
  • The top-k experts are selected based on these scores.
  • The outputs of the selected experts are combined, often using a weighted average or a simple sum.
  • Surprisingly, hashing functions can also be used for routing, achieving some gains even without semantic information.
  • Early attempts to use reinforcement learning (RL) for routing have been largely abandoned due to computational cost and instability.

8. Router Implementation

  • The input token (residual stream) is multiplied by a learned vector for each expert to compute an affinity score.
  • A softmax function normalizes these scores.
  • A top-k function selects the k best experts.
  • The outputs of the selected experts are weighted and summed.
  • The result is added back to the original residual stream.

9. Shared and Fine-Grained Experts

  • Fine-grained experts: Experts with smaller hidden dimensions, allowing for more experts without increasing the parameter count significantly.
  • Shared experts: Experts that are used by all tokens, capturing shared structure in the data.
  • DeepSeek pioneered the use of fine-grained and shared experts.
  • Ablation studies show that fine-grained experts provide significant performance improvements.

10. Common MoE Configurations

  • Early MoE models had large numbers of routed experts.
  • More recent models, like DeepSeek and Quen, use a combination of fine-grained and shared experts.
  • Llama 4 also uses a shared expert.
  • Fine-grained experts are a common feature in modern MoE architectures.

11. Training Challenges and Techniques

  • The main challenge is maintaining sparsity during training while dealing with non-differentiable routing decisions.
  • Reinforcement learning (RL) has been explored but is not widely used due to computational cost.
  • Stochastic approximations, such as adding noise to the router logits, have also been tried but are less common now.
  • The most common approach is to use heuristic loss terms to balance the load across experts.

12. Load Balancing Losses

  • The goal is to prevent a single expert from dominating the training process.
  • A common loss term is the inner product between the fraction of tokens allocated to each expert and the fraction of the router probability allocated to each expert (f.p).
  • This loss penalizes experts that receive a disproportionate number of tokens.
  • DeepSeek uses both per-expert and per-device balancing losses.

13. DeepSeek V3's Auxiliary Loss-Free Balancing

  • DeepSeek V3 introduces a novel approach where they add a fudge factor (b of i) to the softmax scores for each expert.
  • This fudge factor is learned online using a simple gradient scheme, increasing the score for underutilized experts and decreasing it for overutilized experts.
  • However, DeepSeek V3 still uses a complementary sequence-wise auxiliary loss to balance the load at a per-sequence level.

14. Expert and Device Parallelism

  • MoE models are well-suited for expert parallelism, where each expert is placed on a separate device.
  • This requires collective communication calls to route tokens to the appropriate devices and combine the outputs.
  • Modern sparse matrix multiplication engines can optimize computations when multiple experts are on a single device.

15. Token Dropping

  • If an expert receives too many tokens, some tokens may be dropped due to memory limitations.
  • This can introduce stochasticity into the inference process, even with a temperature of zero.

16. Stability Tricks

  • MoE models can be unstable and difficult to fine-tune.
  • Using float32 for router computations and adding an auxiliary z-loss to the router can improve stability.
  • Overfitting can be a concern when fine-tuning MoE models.
  • Using a lot of data or alternating dense and MoE layers can mitigate overfitting.

17. Upcycling

  • Upcycling is a technique to convert a dense model into an MoE model by copying the FFN layers and adding a router.
  • This can be a cost-effective way to obtain an MoE model.

18. DeepSeek V3 Architecture Deep Dive

  • DeepSeek V1 (16B parameters, 2.8B active): Uses a standard top-k routing with shared and fine-grained experts.
  • DeepSeek V2 (236B parameters, 21B active): Architecture is identical to V1, but adds a device selection mechanism to reduce communication costs.
  • DeepSeek V3 (671B parameters, 37B active): Uses auxiliary loss-free balancing and a sequence-wise auxiliary loss.

19. DeepSeek V3 - Multi-Head Latent Attention (MLHA)

  • MLHA is an optimization technique to compress the KV cache.
  • Instead of caching the full hidden state, it projects the hidden state into a lower-dimensional latent space and caches that instead.
  • The projection matrices are merged with the query projection matrix to avoid extra computations.
  • MLHA is not directly compatible with RoPE (Rotary Position Embeddings), requiring a workaround.

20. DeepSeek V3 - Multi-Token Prediction (MTP)

  • MTP is a training technique where the model predicts multiple tokens in parallel.
  • A lightweight one-layer transformer is used to predict the additional tokens.
  • DeepSeek V3 only uses MTP with one token ahead.

21. Conclusion

  • MoE architectures are now central to building high-performance language models.
  • Discrete routing is a key challenge, but heuristic approaches seem to work well in practice.
  • MoE models are cost-effective and worth learning.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.