A Taxonomy for Next-gen Reasoning — Nathan Lambert, Allen Institute (AI2) & Interconnects.ai

AI EngineerAbout 6 min readJul 20, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Reinforcement Learning with Verifiable Rewards (RLVR), Reasoning Models, Inference Time Scaling, Calibration, Strategy, Abstraction, Skills, Tool Use, Overthinking, Planning, Parallel Compute, Continual Learning, Post-Training, Pre-Training.

Main Topics and Key Points

Introduction: Reflecting on Six Months of RLVR

  • The speaker reflects on the progress of reinforcement learning with verifiable rewards (RLVR) six months after the release of models like 01 and DeepSeek.
  • The core idea is that scaling RL at training time leads to improved performance, which then enables inference time scaling.
  • The crucial question is identifying future directions beyond achieving high benchmark scores with large token inputs.
  • The speaker aims to process where the field is heading and what is needed to train models effectively, considering what companies like OpenAI are likely already doing.

Reasoning Models and New Applications

  • Reasoning unlocks new language model applications.
  • Example: 03 directly provides a download link for "coastrunners" after a search query, a task that previously required multiple Google searches.
  • 03 is highlighted as a valuable tool for information retrieval.
  • Deep Research is mentioned as a tool that can be used creatively, such as finding typos on a website.
  • Cloud Code is described as enjoyable to use for less serious software engineering tasks.
  • Codec and fully autonomous agents are emerging, with the potential to be valuable in the near future.
  • The speaker predicts that these tools will become essential for daily use within six months, driven by advancements in reasoning models.

The Expanding Time Horizon of Language Models

  • A plot illustrates the increasing time horizon for task completion by language models.
  • GBT40 showed saturation, but new models in 01 pushed the frontier further.
  • The y-axis represents the approximate time required for models to complete tasks.
  • Reasoning models are key to extending these limits.
  • Continued progress requires human effort to determine what models need to do.
  • Planning is identified as a crucial area for pushing the boundaries and enabling new language modeling applications.

Taxonomy of Traits for Autonomous Reasoning Models

  • The speaker proposes a taxonomy of traits for training autonomous reasoning models:
    • Skills: Proficiency in areas like math and code (largely achieved).
    • Calibration: The ability to align output token usage with problem difficulty (crucial for product development).
    • Strategy: The capacity to choose the right direction and explore different approaches (essential for complex tasks).
    • Abstraction: The capability to autonomously break down problems into tractable subtasks (necessary for very challenging tasks).
  • The speaker emphasizes that models don't natively perform abstraction and need to be trained to do so.
  • Planning is highlighted as a new frontier that requires focused research and implementation.

Reinforcement Learning with Verifiable Rewards (RLVR) Explained

  • RLVR involves prompting an agent, generating completions, scoring them, and updating the model's weights.
  • The process is typically single-turn and relatively simple.
  • The speaker acknowledges the need to update the diagram to reflect multi-turn interactions and tool use.
  • The core remains the same: language models generate completions and receive feedback.

Skill Evaluation and the Need for Planning

  • The speaker references a collection of evals showcasing the progress from GBT40 to 01 and 03.
  • These gains are attributed to the addition of new training methods.
  • The argument is made that a similar approach is needed for planning to succeed.
  • Planning tasks are compared to "humanity's last exam" in terms of difficulty.
  • The list of reasoning abilities and low-level skills is expected to grow, with tool use being a recent addition.
  • 03 is presented as a combination of tool use and reasoning.
  • Abstraction is considered a desirable trait for agentic behavior on top of tool use, but it's difficult to measure.

Calibration and Overthinking

  • Calibration is currently passed to the user through model selectors and reasoning effort selectors, which is not ideal.
  • Models often "overthink," using excessive tokens for simple tasks.
  • Example: Reasoning models may use hundreds or thousands of tokens to answer "2 + 3."
  • Reasoning models can lead to a 10x to 100x increase in token spend.
  • Wasteful token usage can strain infrastructure and increase costs.
  • Users don't want to wait excessively for easy questions or switch models to avoid overthinking.

Strategy and Planning Deficiencies

  • Models currently exhibit limited planning capabilities on their own.
  • Example: When asked to solve a math problem, a DeepSeek model immediately attempts to construct a polynomial without sketching the problem first.
  • This lack of planning can lead to wasted tokens and latency, potentially causing users to abandon the task.
  • Current applications often prompt models to plan manually, but this should be model-native.

Implementation Details for Planning

  • Implementation details for planning include memory management (e.g., Cloud Code compressing memory).
  • The goal is to avoid repeating mistakes and ensure tractable subtasks.
  • Offloading thinking and using parallel compute are suggested for handling challenging parts.
  • Language models should be able to call multiple other models in parallel.

The Effort Behind Reasoning and Planning

  • The development of Qstar (later 01) required significant effort from OpenAI, including 12-18 months of building reasoning traces.
  • A similar effort is needed for planning, but the outputs are more intuitive.
  • Experts can write or check multi-step plans, making hill climbing more feasible.
  • The process involves initial data collection, SFT, and RL to reinforce planning styles.
  • Models can be structured to plan out their answers before thinking.

Skill vs. Planning: A Deeper Dive

  • 03 is highly skilled at search, but lacks planning in applications like Deep Research.
  • This can result in inconsistent outcomes.
  • Improved planning would make models more thorough and reliable.
  • Models can perform searches but struggle to recommend purchases due to a lack of planning.

Taxonomy Revisited and Parallel Compute

  • The four traits (skills, calibration, strategy, abstraction) can be expanded or combined.
  • Parallel compute is highlighted as a way to enhance robustness.
  • 01 Pro remains a strong model, and parallel compute makes it even better.
  • RL training encourages exploration, while parallel compute facilitates exploitation.

Continual Learning and the Future of Training

  • The speaker touches on continual learning and the potential to reduce reliance on pre-training.
  • Scaling up RL is considered a tractable approach.
  • The speaker summarizes their research plan at AI2:
    1. Gather questions with verified answers across various domains.
    2. Filter questions based on difficulty for the base model.
    3. Run stable RL training with continuous improvement.
    4. Apply techniques like overong filtering and reference model resetting.

Post-Training vs. Training

  • The speaker proposes renaming post-training as training.
  • OpenAI's 01 allocated 1% of compute to post-training, while 03 increased it by 10x.
  • Post-training may soon account for a significant portion of compute, potentially matching pre-training.
  • DeepSeek's transition to post-training is noted, with RL training potentially consuming 10-20% of their compute.
  • Scaling RL is a real trend, and it's important to embrace what models can do and break down tasks accordingly.

Conclusion

  • The speaker thanks the audience and invites feedback.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.