Training Agentic Reasoners — Will Brown, Prime Intellect

AI EngineerAbout 4 min readJul 8, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Agentic Reasoners: Agents capable of reasoning and interacting with complex environments using tools.
  • Reinforcement Learning (RL): A machine learning paradigm where an agent learns to make decisions by interacting with an environment to maximize a reward.
  • Reward Hacking: When an agent learns to exploit the reward function to achieve a high score without actually performing the desired task.
  • Rubrics: Fine-grained evaluation criteria used to assess the performance of an agent, which can be generated on the fly.
  • Tools: Capabilities that an agent can use to interact with its environment, such as code execution, web search, or file manipulation.
  • Synthetic Fine-Tuning (SFT): Training a model on data generated synthetically, often used as a warm-up before RL.
  • Proximal Policy Optimization (PPO) & Gradient Regularized Policy Optimization (GRPO): RL algorithms used to train models by iteratively improving the policy while ensuring it doesn't deviate too far from the previous policy.
  • Generator-Verifier Gap: The difference in difficulty between solving a problem and verifying a proposed solution.

RL Works Now

  • DeepSeek Math: A successful example of RL applied at scale, demonstrating that with a good setup, signal, and model, RL can significantly improve performance.
  • OpenAI's Strategy: OpenAI is focusing on scaling RL for future progress, as seen with the 03 release, which prioritizes tool use and agentic capabilities.
  • RL as a Solution for Complex Systems: RL can improve the robustness of systems that tend to fail in complex environments, offering a way to train models to handle a wider range of situations.

Complexity of RL

  • Veril and DeepSeek Architectures: The architectures are complicated with many moving pieces.
  • Need for Accessible Tools: There is a need to simplify RL to make it feasible for startups and individual researchers, bridging the gap between academic research and practical applications.

Agents and RL Are the Same

  • Agentic Products: Products like Cloud Code, Devon, and GPT-4o are successful because the underlying models have been trained using RL to perform specific tasks.
  • RL Training: The process of building an agent (harness, environment, tools, iteration) is conceptually the same as reinforcement learning (policies, actions, states, rewards).
  • Manual RL: Tuning prompts and adjusting harnesses in agent development is akin to performing RL by hand.
  • Key RL Components:
    • Tasks: Prompt variations.
    • Rollouts: Completion sequences.
    • Evaluations: Metrics used to assess performance.
    • Advantage Estimation: Identifying which specific actions or tokens led to better outcomes, allowing for targeted improvements.

RL Algorithms: PPO, GRPO, DPO

  • Proximal Policy Optimization (PPO): Estimates advantages by measuring the difference between expected and actual rewards to refine the model's behavior.
  • Gradient Regularized Policy Optimization (GRPO): A computationally efficient alternative to PPO that leverages sampling to identify advantageous paths.
  • DPO Drawbacks: DPO may lack fine-grained advantage estimation compared to PPO and GRPO.

Avoiding Paper Overload

  • Focus on the Process: It's crucial to understand the holistic process of reinforcement learning and identify the important aspects for solving specific problems.
  • Leverage Software: Focus on practical software and tools that implement RL algorithms effectively, rather than getting bogged down in the details of individual research papers.

The Importance of Tools

  • MCP (Meta-Control Programming): MCP is fundamentally about providing LMs with tools to interact with environments and solve problems.
  • Real-World Tasks: RL should focus on real-world tasks and the challenges that arise when designing reward functions for these tasks.

Reward Hacking and Good Evals

  • Reward Hacking: A real concern where models exploit the reward function instead of learning the intended behavior.
  • Building Good Evals: Creating reward signals that accurately capture the desired behavior and make it more difficult to "game" the system than to perform the task correctly.
  • Generator-Verifier Setup:
    • Evaluation on ambiguous tasks can be done by breaking down tasks into smaller pieces.
    • Using LMS as subroutines in evaluations (LM judge on steroids)
    • Train a specialized LM to do fine-grain evaluations.

Multi-Turn Agentic Systems

  • Future of RL: Multi-turn interactions are essential for tasks like agentic search, tool use, and long-horizon planning.
  • Conceptual Pieces:
    • Environments: Harnesses.
    • Rewards: Evals.
    • Tasks: Prompts.
    • Policy: LM API

Verifiers Library

  • Toolkit for RL: A toolkit designed to simplify the process of building and training agents with RL, making it feel like standard agent development.
  • Interface: A client object compatible with OpenAI-like APIs.
  • Components: Parsers and rubrics.
  • SF Warm-up: Using synthetic data loops or evals with APIs like Claude or OpenAI to debug and warm-up models before RL.
  • Efficiency: Asynchronous computation to effectively utilize resources.

Conclusion

The future of agentic software relies on effectively integrating reinforcement learning to create robust and capable agents. While RL can be complex, tools like the Verifiers library are emerging to simplify the process, making it accessible to a wider audience. The key is to focus on building good evaluation metrics (rubrics) that incentivize desired behaviors and to embrace multi-turn interactions for complex tasks. With the right approach, RL can be a powerful tool for creating agents that can reason, adapt, and solve real-world problems.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.