Key Concepts
- GPT OSS: OpenAI's open-weights mixture of experts model.
- Mixture of Experts (MoE): A model architecture where only a subset of parameters (experts) are activated for each token.
- Grouped Query Attention (GQA): An attention mechanism that reduces memory usage by sharing key-value pairs across multiple query heads.
- SwiGLU: An activation function used in feed-forward network layers.
- Rotary Positional Embeddings (RoPE): A method to encode token position directly into the attention mechanism.
- RMSNorm: A normalization method that scales inputs by their root mean square.
- Yarn Scaling: A technique to extend the context window of a model by adjusting the frequency of RoPE.
- Quantization: Reducing the precision of model weights to decrease size and improve inference speed.
- Quen 3: Alibaba Cloud's family of open-source models, including dense and MoE variants.
- QK Norm: A normalization step that dynamically rescales query and key vectors.
- ABF (Attention Bias Fine-tuning): A technique to adjust RoPE for longer sequences.
- Dual Chunk Attention: A method to process long sequences efficiently.
- Long Chain of Thought Cold Start: A post-training stage involving feeding the model challenging reasoning problems.
- GRPO: An RL algorithm used for reasoning.
- Thinking Mode Fusion: Integrating reasoning and non-reasoning capabilities into a single model.
- Strong to Weak Distillation: Training smaller models from larger ones.
- DeepSeek V3: A large open-source MoE model.
- MLA (Multi-head Latent Attention): An attention mechanism that compresses keys and values into a smaller latent space.
GPT OSS: OpenAI's Open-Weights Model
- Overview: GPT OSS is OpenAI's first major open-weights model since 2019, available in 120 billion and 20 billion parameter sizes. It's a mixture of experts model, where only the top four experts are activated per token.
- Architecture:
- Decoder-only transformer architecture.
- Employs Grouped Query Attention (GQA) for memory efficiency.
- Uses SwiGLU activations in feed-forward layers.
- Incorporates Rotary Positional Embeddings (RoPE) for encoding token position.
- Utilizes RMSNorm with pre-normalization for stable training.
- Context Window: Achieves a 131,000 token context window using Yarn scaling during pre-training.
- Tokenizer: Uses OpenAI's open-source O200K harmony tokenizer, a byte pair encoding tokenizer with over 200,000 tokens.
- Training Data: Trained on a text-only corpus in the trillions of tokens, focusing on STEM coding and general knowledge. Harmful content was filtered out.
- Post-Training: Underwent substantial post-training for safety and alignment. Released in a quantized format by default.
Quen 3: Alibaba Cloud's Open-Source LLM
- Overview: Quen 3 is Alibaba Cloud's family of models, including both dense and mixture of expert variants. Dense models range from 6 billion to 32 billion parameters, while MoE models come in two sizes.
- Architecture:
- Dense models are similar to Quen 2.5, incorporating GQA, SwiGLU, RoPE, and RMSNorm.
- MoE models share the same fundamental architecture as dense models but add a mixture of experts layer with 128 total experts, activating eight per token.
- Uses a byte-level byte pair encoding tokenizer.
- QK Norm: Replaces QKV bias with QK Norm, a normalization step that dynamically rescales query and key vectors.
- Training Data: Trained on 36 trillion pre-training tokens, including multilingual texts, STEM and coding sources, and reasoning tasks. Uses Quen 2.5 models to generate synthetic data.
- Training Stages:
- Stage 1 (General): Trained on over 30 trillion tokens covering 119 languages at a sequence length of 4096 tokens.
- Stage 2 (Reasoning): Trained on an additional 5 trillion higher quality tokens featuring more STEM reasoning and coding problems.
- Stage 3 (Long Context): Context length extended to over 32,000 tokens using ABF, Yarn, and dual chunk attention.
- Post-Training Pipeline:
- Long Chain of Thought Cold Start: Trained on challenging reasoning problems with verifiable answers.
- Reasoning RL: Uses GRPO on roughly 4,000 query-verifier pairs.
- Thinking Mode Fusion: Integrates reasoning and non-reasoning into a single model.
- General RL: Broadens capabilities in instruction following, formatting, preference alignment, tool use, and specialized scenarios.
- Strong to Weak Distillation: Training of smaller models from larger ones.
DeepSeek V3: A Large Mixture of Experts Model
- Overview: DeepSeek V3 is a 671 billion parameter mixture of experts model designed for efficiency and capability.
- Architecture: Mixture of experts model with hardware and algorithmic optimizations, including training natively in 8-bit.
- V3.1 Update: Builds on the original V3 checkpoint, extending it with a two-phase long context training approach and adding a hybrid thinking mode. Improves tool use and agent performance.
- MLA: Uses Multi-head Latent Attention (MLA), which compresses keys and values into a smaller latent space before caching them.
Model Comparison: Size, Context Length, and Techniques
- Size:
- Quen 3: Offers both dense (6B-32B) and MoE (30B, 235B) variants.
- DeepSeek V3: MoE architecture with 671B parameters (37B active).
- GPT OSS: 117B (5.1B active) and 21B (3.6B active) parameter models.
- Context Length:
- GPT OSS: Applies Yarn during pre-training for a native 131,000 token context.
- DeepSeek V3: Fine-tunes to 32,000 tokens, then further trains to 128,000 tokens.
- Quen 3: Fine-tunes to 32,000, then applies Yarn scaling at inference time to reach 128,000 tokens.
Key Takeaways and Observations
- Empirical Findings: Many advancements in deep learning are based on empirical findings rather than first-principles justifications.
- Similar Results, Different Techniques: Models with similar benchmark statistics often achieve these results using different techniques.
- Reinforcement Learning: Reinforcement learning is heavily used in post-training and reasoning, sometimes requiring surprisingly little data.
- Dataset Engineering: The differences in datasets between labs are significant and likely contribute to the unique capabilities of each model.
- Focus on Methods: When evaluating models, focus on the specific methods used to achieve results rather than just benchmark performance.
- Opaque Data Differences: It's difficult to discern the specific differences in datasets used by different labs. This data engineering is likely a significant competitive advantage.
Synthesis/Conclusion
The landscape of open-source LLMs is rapidly evolving, with models like GPT OSS, Quen 3, and DeepSeek V3 pushing the boundaries of size, context length, and performance. While these models share common architectural elements, they differ significantly in their training techniques, data sets, and specific optimizations. Understanding these nuances is crucial for effectively utilizing and further developing these powerful tools. The focus should be on the specific methods employed by each lab rather than solely relying on topline benchmark statistics.
AI summaries can miss context or contain errors. Check important details against the original video.





