Key Concepts
- Paper Club: A weekly event to discuss AI research papers.
- Test of Time Paper Club: A new, curriculum-based version of Paper Club focusing on foundational AI papers.
- DeepSeek V3: A large language model known for its reasoning capabilities.
- DeepSeek R1 (May 28th Update): An improved version of DeepSeek V3 with enhanced reasoning, function calling, and JSON output.
- Quen 3 8B (Distilled): A smaller language model distilled from DeepSeek R1, achieving impressive performance.
- Test Time Scaling: Improving model performance by increasing computation during inference (reasoning).
- gRPO (Generalized Proximal Policy Optimization): A reinforcement learning algorithm used to train DeepSeek models.
- SFT (Supervised Fine-Tuning): Fine-tuning a model on a labeled dataset.
- RLHF (Reinforcement Learning from Human Feedback): Training a model using human preferences as a reward signal.
- Distillation: Training a smaller model to mimic the behavior of a larger, more powerful model.
- Aha Moments: Instances where a model demonstrates a sudden insight or understanding during reasoning.
Paper Club Year in Review and Test of Time Paper Club Launch
- Paper Club has been running weekly for a year and a half, exceeding initial expectations.
- Features authors from Nvidia, Meta, Allen AI, and Amazon, providing direct interaction and feedback.
- Average attendance of 100 people on Wednesdays at noon, with DeepSeek V3 attracting 300 live participants.
- Built on volunteer efforts, fostering a community of AI enthusiasts.
- Launching "Test of Time Paper Club" as a V2, curriculum-based version alongside the original.
- Original Paper Club will continue weekly, covering trending papers with author Q&A.
- Test of Time Paper Club will focus on core AI engineering knowledge, covering fundamental papers.
- Curriculum will cover themes like attention, sequential text generation (GPT-2), optimizers, and key inference techniques (speculative decoding, flash attention, stable diffusion, whisper).
- Kicking off in July and running until December (six months).
- Each week will cover two to four pre-presented papers, allowing for in-depth exploration.
- Aiming to cover 50-100 papers in six months, focusing on core concepts.
- Each week will focus on a different topic, such as whisper (speech-to-text) or image generation (CLIP, stable diffusion).
- Introducing in-person sessions in San Francisco, alongside the remote option.
Test of Time Paper Club: Topics and Structure
- Topics are still open for discussion and input from the community.
- Proposed categories include:
- Foundations of Deep Learning: Attention, optimization, ReLU, gradient descent, basic RL.
- LLM Foundations: RNNs, LSTMs, bidirectional RNNs, BERT, GPT-2.
- Generative LLMs: Llama 3, DeepSeek, other core LLMs.
- Pre-training and Post-training: Scaling laws, Chinchilla, distillation, small model scaling laws (e.g., Phi team).
- Generative Models: CLIP, Sora, Segment Anything, diffusion.
- Fine-tuning: LoRA, QLoRA, DPO, RLGRPO.
- Voice: Whisper.
- Optimization: Speculative decoding, flash attention.
- Eval Tracks: Rexus, Eugene Yan will contribute.
- Each topic will be covered in one to two sessions, with three to four papers presented per week.
- Volunteers are encouraged to present papers.
- SF sessions will have a venue, while remote sessions will continue via Zoom.
- Goal is to provide in-depth understanding rather than superficial overviews.
- Discord channel for discussion and topic suggestions.
- Google Form for recommending papers and volunteering as a speaker.
- Curriculum-based approach allows for a structured learning experience.
- Sessions will be recorded and available on YouTube for later viewing.
DeepSeek V3 and R1: A Deep Dive
- Today's paper club focuses on DeepSeek, specifically the May 28th update.
- DeepSeek V3 was a popular paper, but many haven't had time to review it.
- The May 28th update is a significant improvement, despite being labeled as a minor revision.
- Simon Willis mentioned the lack of good naming conventions for models, highlighting the substantial improvements in DeepSeek R1.
- The update primarily involves better post-training on DeepSeek V3, resulting in enhanced reasoning capabilities.
- AIM 2024 score improved from 70% to 87.5%, matching 03 and 2.5 level performance in math, coding, and reasoning.
- The original DeepSeek V3 required 12,000 tokens to reason through the AIM benchmark, while the updated model requires 25,000 tokens, indicating improved reasoning ability.
- This demonstrates the potential of scaling in the "test time compute" dimension.
- The updated model is also better at JSON output and function calling.
- Benchmarks show that the new DeepSeek R1 is comparable to 03 and Gemini 2.5.
DeepSeek V3 and R1: Distillation
- DeepSeek also released a new distillation model based on Quen 3 8B.
- This new distillation model outperforms the previous distillation model by 10%.
- The Quen 3 8B distilled model matches the performance of the Quen 3 235B "thinking" model.
- This highlights the effectiveness of distilling from a better reasoning model.
- Chain-of-thought improvements distill down effectively, leading to significant performance gains in smaller models.
DeepSeek V3 and R1: Original Paper Recap
- DeepSeek V3 was the first test-time scaling open model.
- Two models were released: DeepSeek R10 and DeepSeek R1.
- DeepSeek R10 is a base model trained with gRPO RL, exhibiting emergence capabilities, reflection, and aha moments.
- DeepSeek R1 is a four-stage pipeline involving cold start reasoning, RL, rejection sampling, and SFT.
- The goal was to explore the potential of LLMs to develop reasoning capabilities without supervised data, focusing on self-evolution through pure RL.
- The DeepSeek team post-trained the DeepSeek V3 base with gRPO, observing the emergence of reasoning abilities.
- The four-step approach to training R1 involves:
- Cold start with SFT.
- RL for reasoning.
- Rejection sampling for generation purposes.
- RL polishing.
- The shift from next-token predictors to scaling in another access (reasoning) was crucial.
- Instead of spending the same compute for every token, the goal was to dynamically spend more compute on different queries.
- Models were trained with pure RL on verifiable outputs (code and math).
- This led to the emergence of reflection and aha moments.
- DeepSeek showed that RL can be done from base models to achieve reasoning, and that distillation is effective for small models.
DeepSeek V3 and R1: Inference Time Scaling
- Inference time scaling involves increasing the chain-of-thought reasoning process, allowing models to spend more time thinking before responding.
- This is an alternative to scaling up pre-training, which becomes exponentially more expensive.
- DeepSeek used pure RL to achieve reasoning capabilities without supervised data.
- DeepSeek V3 (the precursor to R1) had 37 billion active parameters and was trained on 15 trillion tokens.
- DeepSeek V3 introduced multi-headed latent attention.
- The DeepSeek team was able to catch up to the United States due to constraints (trade restrictions), forcing them to be clever with limited GPUs.
- They focused on inference optimization and RL.
DeepSeek V3 and R1: R10 vs. R1
- DeepSeek R10 is a reasoning model trained only on unraveled chain of thought with RL, but it's not a great general model.
- DeepSeek R1 is trained using the outputs of R10, using the four stages of training.
- DeepSeek R10 uses gRPO for RL, with rewards based on accuracy and format.
- The model is trained to output its thinking process.
- DeepSeek R10 exhibits the ability to solve complex tasks by extending test time compute.
- As test time compute increases, interesting behaviors emerge, such as reflection and aha moments.
- Reflection involves models revisiting and re-evaluating previous steps.
- Aha moments are instances where the model demonstrates a sudden insight or understanding.
- To make DeepSeek R1 a useful assistant, the following steps were taken:
- Cold start with strong SFT to prevent instability.
- RL on hard code and math problems.
- Rejection sampling on examples that don't work.
- Final RL stage for general use.
- The final RL stage aims to make the model helpful, harmless, and a good reasoner.
DeepSeek V3 and R1: Performance and Distillation Details
- The new DeepSeek R1 (May 28th) is significantly better than the original, with improved reasoning, function calling, and JSON output.
- It is now comparable to 03 and Gemini 2.5.
- The AIM score jumped 17.5%, and reasoning tokens doubled.
- The new distillation model (Quen 3 8B) is also a significant improvement.
- Distillation involves SFT-style training on reasoning traces.
- Distillation models outperform their base models.
- RL can further improve performance, but it's important to have a good base model.
- Future work includes improving function calling, multi-turn capabilities, and language mixing.
Synthesis/Conclusion
The presentation highlights the evolution of the Paper Club and the launch of the Test of Time Paper Club, emphasizing the importance of foundational knowledge in AI. It also provides a detailed analysis of the DeepSeek models, particularly the May 28th update, showcasing the significant improvements in reasoning capabilities and the effectiveness of distillation techniques. The key takeaways are the power of test time scaling, the importance of RL in achieving reasoning abilities, and the potential of distillation for creating high-performing small models. The speaker encourages community involvement in shaping the curriculum of the Test of Time Paper Club and emphasizes the collaborative nature of the Paper Club initiative.
AI summaries can miss context or contain errors. Check important details against the original video.