How do thinking and reasoning models work?

By Google for Developers

Share:

Key Concepts

  • Thinking Models/Reasoning Models: Large Language Models (LLMs) designed to perform complex tasks by generating intermediate reasoning steps.
  • Scaling Laws: The principle that model performance improves with increased training data, compute, and parameters.
  • Test Time Compute: The application of additional computational resources during the inference (response generation) phase to improve model performance.
  • Chain of Thought (CoT) Prompting: A technique where LLMs are prompted to generate a series of intermediate steps leading to a final answer, improving accuracy on reasoning tasks.
  • Best of N: A test-time compute strategy where multiple responses are generated for a single prompt, and the most frequent response is selected.
  • Learned Reward Models: Models used to evaluate and score candidate responses based on specific criteria, aiding in selecting the best answer.
  • Reinforcement Learning (RL): A machine learning paradigm where an agent learns to perform a task by interacting with an environment, receiving rewards or penalties.
  • Supervised Fine-Tuning (SFT): A post-training technique where models are trained on input-output pairs to learn specific behaviors or formats.
  • Verifiable Reward: A reward signal in RL that is objectively determined, such as a correct answer to a math problem.

How Thinking Models Work

This video explains the mechanisms behind "thinking models" or "reasoning models," a type of LLM that excels at complex tasks by generating intermediate reasoning steps.

1. Scaling Laws and Test Time Compute

  • Scaling Laws: The foundational principle is that LLM performance scales with increased training data, compute, and model parameters. This has historically led to better text generation, problem-solving, and coding capabilities.
  • Test Time Compute: Researchers explored whether increasing compute during inference (when the model generates a response) could also improve performance. This concept is known as "test time compute."

2. Chain of Thought Prompting

  • Concept: Chain of Thought (CoT) prompting is a method to enhance LLM reasoning by instructing the model to generate a series of intermediate steps before providing the final answer.
  • Mechanism: LLMs generate responses token by token, with each token requiring a forward pass through the model's weights. Without CoT, the model must solve a complex problem in a single pass. With CoT, breaking down the problem into smaller steps allows for multiple forward passes, effectively dedicating more compute to reasoning.
  • Example: In a math problem, a model without CoT might incorrectly state "The cafeteria does not have 27 apples." However, when prompted to show its work (e.g., "The cafeteria has nine apples."), it correctly solves the problem. This demonstrates that encouraging the model to "show its work" leads to more accurate results.

3. Strategies for Utilizing Test Time Compute

  • Best of N:
    • Process: Generate 'n' (e.g., 100) responses to a single prompt using the same model, potentially varying diversity with parameters like temperature.
    • Selection: The response that appears most frequently among the 'n' generated outputs is chosen.
    • Limitations: While increasing compute, this method can plateau in performance for nuanced tasks where errors are consistently repeated across reasoning steps. It's likened to repeatedly baking a cake with mislabeled ingredients – more attempts won't fix the fundamental error.
  • Learned Reward Models:
    • Process: A second model, a "reward model," is used to score each of the 'n' candidate responses.
    • Scoring: The reward model assigns a score based on criteria like correctness, fluency, or relevance. For the math problem, a well-calibrated reward model would score the correct answer (9) highly and the incorrect answer (27) lowly.
    • Selection: The response with the highest score is returned. Variations include grouping identical answers and selecting from the bucket with the highest cumulative score, or using a reward model that scores each step of the chain of thought.

4. Reinforcement Learning for Enhanced Reasoning

  • Role of RL: Reinforcement Learning (RL) is crucial for teaching models to produce longer chains of thought during post-training.
  • RL Framework:
    • Agent: The LLM being tuned.
    • Environment: The context window, including the prompt and generated text.
    • Actions: Generating tokens.
    • State: The current content of the context window.
    • Reward: Feedback on the agent's actions.
  • Application to LLMs: For reasoning tasks, the LLM attempts problems with objective answers (e.g., math, logic). The reward is "verifiable," meaning it's a direct check of correctness. This provides unambiguous feedback for the model to learn.
  • Learning Process: Unlike supervised learning, the LLM isn't shown explicit input-output pairs. Instead, it learns by exploring possible token sequences, receiving rewards for correct answers, and adjusting its parameters to maximize future rewards.
  • Emergent Behavior: A key finding is that RL training often leads to longer reasoning chains, which correlate with improved performance.

5. Combining Supervised Fine-Tuning and Reinforcement Learning

  • Synergy: Research suggests that a combination of Supervised Fine-Tuning (SFT) and RL is highly effective.
  • SFT's Role: SFT bootstraps the RL process by teaching the model to follow instructions and produce answers in a consistent format.
  • RL's Role: RL then helps the model generalize its reasoning capabilities, learning underlying rules and knowledge applicable to new, unseen problems. The paper "SFT memorizes, RL generalizes" highlights this.

6. Synthesis and Conclusion

  • Core Idea: Thinking models work by effectively utilizing compute at test time.
  • Mechanisms: This is achieved through techniques like Chain of Thought prompting, which allows for more intermediate reasoning steps, and post-training methods like Reinforcement Learning, which trains models to generate longer and more effective reasoning chains.
  • Deployment Strategies: After training, strategies like "Best of N" or using reward models can further refine the selection of the best response.
  • Future: The field is rapidly evolving, with ongoing research into optimal methodologies for state-of-the-art models.

The video concludes by encouraging viewers to explore Gemini's thinking capabilities and links to relevant research papers for deeper understanding.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video