Deepseek R1 - The Era of Reasoning models

AI JasonAbout 5 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Reasoning Models (e.g., DeepSeek R1, OpenAI's O1): AI models designed for extended, high-quality reasoning.
  • Chain of Thought (CoT): A prompting technique where the model is encouraged to think step-by-step.
  • Knowledge Distillation: Training a smaller model using data generated by a larger, more capable model.
  • Zero-Shot Prompting: Providing only the task instruction without any examples.
  • Few-Shot Prompting: Providing a limited number of examples (one or two) along with the task instruction.
  • Agent Planning: Using reasoning models to generate detailed plans for agents to execute complex tasks.
  • Reinforcement Learning: A training method used to incentivize models to generate higher quality and longer reasoning tokens.

DeepSeek R1 and the Rise of Reasoning Models

DeepSeek released R1, an open-source reasoning model, that the presenter claims has comparable performance to OpenAI's O1 but is significantly cheaper (96%). The DeepSeek team has also revealed the training methods used for this model. A distilled version of R1 can be run on a home device with a CPU and at least 48 GB of RAM. The presenter believes that reasoning models will be a key driver of advancements in large language models (LLMs) in 2025.

The Importance of Reasoning and Computation

One key finding from OpenAI's O1 model is that the longer the model "thinks," the better the results. This suggests that scaling model capabilities can be achieved through increased computation during the inference stage, rather than solely relying on pre-training data. The presenter references Ilya Sutskever's statement that "pre-training as we know it will end," implying that the future of LLM advancement lies in enhancing reasoning capabilities during inference.

How Reasoning Models Work

Reasoning models build upon the "chain of thought" technique by generating longer and higher-quality reasoning chains. These models exhibit behaviors such as:

  • Stopping and re-evaluating their approach.
  • Breaking down problems into smaller steps.
  • Trying multiple different strategies.

These behaviors are not explicitly programmed but emerge through reinforcement learning, where the model is incentivized to generate higher-quality and longer reasoning tokens. The DeepSeek R1 paper highlights the "emergence of sophisticated behavior such as reflection" and "exploration of alternative approaches to problem-solving" as spontaneous developments.

Knowledge Distillation with Reasoning Models

DeepSeek's R1 model can be used to generate high-quality reasoning data, which can then be used to train smaller models for specific domain tasks through knowledge distillation. This allows for good performance on edge devices. The open-source nature of R1 is significant because OpenAI's O1 model has kept its reasoning tokens hidden, preventing developers from using them for fine-tuning or knowledge distillation. The presenter believes knowledge distillation will be a significant trend in 2025.

Prompting Best Practices for Reasoning Models

Despite their power, reasoning models come with trade-offs in terms of cost and latency. The presenter outlines several best practices for prompting reasoning models effectively:

1. Keep Prompts Simple and Direct

Unlike older models, complex prompting techniques often degrade performance with reasoning models. A paper titled "Do Advanced Large Language Models Eliminate the Need for Prompt Engineering?" found that zero-shot prompting (simply stating the task) often yields the best results with reasoning models. Open AI suggests avoiding detailed instructions and prefixes that guide the response generation.

2. One- to Two-Shot Prompting

While few-shot prompting can be useful, providing too many examples (more than two) can decrease performance. The presenter references a paper called MacQA, which found that O1's output was lower with five examples compared to minimal prompting. Open AI recommends showing the model how to do it with a limited number of relevant examples to prevent overcomplicating the response.

3. Prompt for Extended Reasoning

Encouraging the model to "take your time and think carefully" can lead to more reasoning tokens and increased accuracy. Research has shown that prompting for extended reasoning can increase reasoning tokens by 16-30% and improve accuracy.

When to Use Reasoning Models

Reasoning models should not replace day-to-day models but offer a new option for increased intelligence at the cost of higher latency and cost. The presenter suggests decomposing tasks into smaller steps and identifying those that would benefit from additional reasoning, where latency is less critical.

Use Cases:

  • Agent Planning and Reasoning: Reasoning models can generate plans for agents to execute complex tasks, which can then be passed on to smaller models for execution.
  • Image Reasoning and Understanding: Reasoning models excel at understanding complex images like flowcharts and diagrams, making them useful for medical or image pre-processing tasks.

Example: Agent Planning for Logistics

The presenter provides an example of using O1 to generate a plan for a logistics agent to determine the best route to fulfill a customer order. The O1 model generates a detailed, step-by-step plan with clear "if-then-else" statements, which is then executed by a smaller, cheaper model like Full Mini. This demonstrates how reasoning models can drive complex decision-making actions that smaller models cannot handle alone.

Conclusion

Reasoning models like DeepSeek R1 and OpenAI's O1 represent a significant advancement in LLM capabilities. By understanding how these models work and employing effective prompting techniques, developers can leverage their power for complex tasks such as agent planning and image understanding. While there are trade-offs in terms of cost and latency, the potential benefits of increased intelligence and reasoning capabilities make these models a valuable tool for building sophisticated AI applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.