Bringing new tool use advancements to life: Claude Plays Pokemon

AnthropicAbout 5 min readAug 1, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Tool Use: Utilizing external functions or APIs to enhance model capabilities.
  • Agentic Loop: A cycle of planning, acting, learning, and repeating to achieve goals.
  • Extended Thinking Mode: Allowing the model to deliberate and reflect between tool calls.
  • Parallel Tool Calling: Executing multiple tool calls simultaneously for efficiency.
  • Tool Design: Structuring and describing tools effectively for model understanding.
  • Prompt Engineering: Crafting precise instructions to guide model behavior.

1. Introduction: QuadPlays Pokemon and Tool Use

  • The speaker, David, introduces QuadPlays Pokemon as a demonstration of improved tool use in new models.
  • The focus shifts from a general discussion of tool use to a specific example of a model playing Pokemon.
  • The relaunch of QuadPlays Pokemon is presented as a live demonstration of these improvements.

2. Improvements in Tool Calling

  • Extended Thinking Mode:
    • Models can now build plans, reflect, and question assumptions between tool calls.
    • Example: In the name entry screen, the model can now recognize and correct cursor errors by analyzing its previous actions and adjusting its strategy.
    • Quote: "I'm really stumped. The cursor went right instead of left when I hit left. What happened?" - Example of model self-reflection.
  • Parallel Tool Calling:
    • Models can now call multiple tools simultaneously, improving efficiency.
    • Example: The model can now talk to mom and update its knowledge base at the same time, saving time and tokens.
    • In Quad 3.7, the model would call one tool and wait, requiring a new generation for each subsequent action.
    • This reduces the "plan-act loop" frequency.

3. Evolution of Tool Use

  • Initially, tool use was primarily for tasks like math calculations (using a calculator tool).
  • Now, tool use is the foundation for building agents capable of long-term, complex actions.
  • Agentic Loop: Plan -> Act -> Learn -> Repeat.
    • Example: In Pokemon, the model plans to talk to mom, presses A, reflects on the dialogue box, and continues playing.

4. Interleaved Thinking: A Deeper Dive

  • In previous versions, the model would create a complete, often flawed, plan at the beginning.
  • Example: The model would plan to enter its name and beat Pokemon, but would fail at the name entry screen due to cursor errors.
  • With extended thinking, the model can now adapt its plan based on real-time feedback.
  • Example: The model recognizes that the cursor wrapped around the screen and adjusts its strategy accordingly.

5. Parallel Tool Calling: Efficiency Gains

  • Parallel tool calling saves time by allowing the model to perform multiple actions simultaneously.
  • Example: The model can talk to mom, update its knowledge base, and advance the dialogue multiple times in a single step.
  • This reduces redundant tool calls and improves agent speed.

6. Training Models as Agents

  • Anthropic focuses on making models smarter for long-term, complex problem-solving.
  • Models are trained to be useful and easy to build with, incorporating developer feedback.
  • Parallel tool calling is an example of a feature developed based on user needs.

7. Q&A: Tool Design and Implementation

  • Hierarchy of Tools:
    • The speaker emphasizes the importance of clear tool design and separation of concerns.
    • Example: A tool for navigating the overworld vs. pressing buttons in battle.
    • The key is to provide clear descriptions and examples of when to use each tool.
  • Tool Descriptions vs. Prompts:
    • The speaker suggests that information can be placed in either the tool description or the prompt.
    • Tool descriptions ensure the model understands the syntax.
    • Clear descriptions are more important than the specific location.
  • Number of Tools:
    • Models can handle 50-100 tools.
    • The challenge is in defining the tools precisely and avoiding overlap.
  • Helper Functions/Proxies:
    • Using helper functions to manage a large number of tools.
    • The speaker believes that smarter models can handle full context and complex decisions.

8. Claude Performance in Pokemon

  • Opus is significantly better at Pokemon, but not necessarily in visually understanding the Game Boy screen.
  • Its planning and execution abilities are greatly improved.
  • Example: The model successfully completed a 24-hour grind session to catch 10 Pokemon and obtain Flash.
  • The model may still get stuck in Mount Moon, but it will do more intelligent things in the process.

9. Parallel Tool Calling: State of the Art

  • Parallel tool calling is not necessarily state-of-the-art, but it is a useful feature.
  • The model now understands that it can make multiple tool calls at once.
  • The API returns an object with multiple tool use blocks.

10. Potential Issues with Planning

  • The model may sometimes be too impatient and restart conversations by spamming the A button.
  • This is due to a lack of understanding of time and the inability to see the intermediate steps.
  • This can be addressed with prompting, helping the model understand its limitations.
  • Example: Telling the model that it won't see the results between button presses.

11. Consistency with Multiple Tools

  • The new models are better at precise instruction following.
  • Good tool design and crisp prompting are key to achieving consistent performance across many tools.
  • There is more room to improve prompts and achieve better results with a wider range of tools.

12. Conclusion

  • The speaker acknowledges that the talk deviated from the original plan.
  • The key takeaway is that models are getting better at being agents through improvements in tool use, planning, and instruction following.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.