Key Concepts
- Open-source reasoning data sets
- Supervised Fine-Tuning (SFT)
- Reinforcement Learning (RL)
- Reasoning data recipe
- Teacher models
- Distillation
- Data set pipeline: sourcing, mixing, filtering questions, generating answers, filtering answers
- Scaling laws
- Context window
- Verification
- Evaluation
Open Thoughts: Creating Open-Source Reasoning Data Sets
Ryan from Bespoke Labs discusses Open Thoughts, a project focused on creating high-quality open-source reasoning data sets. The presentation emphasizes the importance of reasoning in AI models and the surprising effectiveness of SFT in achieving strong reasoning capabilities, particularly highlighted by DeepSeek R1.
The Rise of Reasoning in AI
- Recent advancements show significant performance improvements in models on reasoning benchmarks like Amy (competitive math problems).
- Models benefit from step-by-step thinking, enabling long chains of thought.
- DeepSeek R1, an SFT model fine-tuned on 800K examples (600K reasoning), demonstrates the power of SFT, even with RL used for data generation and alignment.
- The success of DeepSeek's small reasoning models motivated the Open Thoughts project.
The Missing Link: The Reasoning Data Recipe
- While the training recipe (SFT) is known, the data recipe for creating strong reasoning models is not well-defined.
- Key questions include: How much data is needed? What data creation steps are necessary? What are the optimal choices for each step?
Why Train Your Own Reasoning Models?
- Performance, privacy, speed, cost, ownership, and destiny are key reasons.
- Reasoning is a valuable tool for solving domain-specific problems.
- SFT is presented as an easy and effective approach.
Open Thoughts 3: A State-of-the-Art Reasoning Data Set Recipe
- Open Thoughts 3, the latest version, aims to provide a state-of-the-art recipe for reasoning data sets.
- Benchmarks: Amy (competitive math), Live Codebench (competitive code), and GPQA Diamond (science questions).
- Results show improved accuracy with increased data scale.
- Open Thoughts 3 outperforms the Neimatron nano data set (Nvidia) when training on the same base model.
- The project surpasses the Deepseek R1 quen 7B model in performance.
Achieving State-of-the-Art Performance
- Scaling data set size is a major factor in increasing accuracy, but it becomes exponentially more expensive.
- Improving the data set recipe shifts the scaling curve upwards.
The Data Set Pipeline
- The pipeline is broken down into:
- Sourcing questions: Gathering raw questions from various sources.
- Mixing: Combining different sources of questions.
- Filtering questions: Selecting high-quality questions.
- Generating answers: Using a teacher model for distillation.
- Filtering answers: Removing incorrect or low-quality answers.
- Selecting teacher models: Determining the best model for answer generation.
- The project involved over 5,000 data sets and almost 3,000 models, with around a thousand experiments specifically for this project.
Key Learnings from the Data Set Recipe
- Sampling Multiple Answers: Sampling multiple reasoning traces per question significantly improves performance. Scaling by 16x is possible by sampling 16 times per question.
- Teacher Model Selection: A model's performance on evaluation benchmarks does not necessarily correlate with its effectiveness as a teacher model. Quen 32B was found to be a better teacher than Deepseek R1.
- Synthetic Questions: Data sources with synthetic questions performed well, offering scalability.
- Question Filtering: Filtering questions based on difficulty (using language models) and response length (proxy for difficulty) proved effective.
- Diversity: Choosing a smaller number of high-quality sources is better than optimizing for diversity with a larger number of sources.
- Verification: Filtering based on answer verification did not significantly improve performance for SFT and distillation.
Adapting the Recipe for Specialized Reasoning Models
- Be aware that the optimal choices may vary based on the domain.
- Start with the Open Thoughts recipe and iterate.
- Study each step in the pipeline distinctly for different domains (e.g., code, science, math).
- Utilize synthetic question generation to expand data sets, especially when limited data is available.
- Use the open-source library "curator" for this.
- Evaluation is paramount. Use the open-source library "Evalchemy" for evaluation, sharding, and parallelism. For small evaluation sets, run the model multiple times and average the results.
Surpassing the Teacher with Distillation
- Distillation can surpass the teacher model in some domains.
- In legal reasoning (classifying Supreme Court decisions), a fine-tuned 7B model surpassed R1 by using 2k unique questions, sampling five answers per question, and verifying the answers.
Open Source Resources
- Paper, weights, data set, and code repositories are available.
Conclusion
The Open Thoughts project provides valuable insights and resources for creating high-quality reasoning data sets. The key takeaways include the importance of multiple reasoning traces, careful teacher model selection, the effectiveness of synthetic data, and the potential to surpass teacher models through distillation. The open-source nature of the project encourages further research and application in various domains.
AI summaries can miss context or contain errors. Check important details against the original video.