Key Concepts
- Advanced Tool Use (Anthropic): A new suite of features in Claude designed to optimize agent interaction with tools.
- Tool Search Tool: A mechanism allowing agents to discover and load relevant tool definitions on demand, rather than loading all tools upfront.
- Programmatic Tool Calling: An approach where agents write and execute code (e.g., Python scripts) to orchestrate multiple tool calls and process intermediate results, only exposing the final output to the agent's context.
- Tool Use Examples: Small, realistic examples of tool calls embedded directly into tool definitions to guide the model on proper usage.
- MCP (Multi-tool Code Protocol/Server): A standard or framework for connecting agents to multiple tools, often leading to high token consumption in its traditional implementation.
- Context Window: The limited amount of information an LLM can process at one time, measured in tokens.
- Token Savings: Reduction in the number of tokens consumed, leading to cost efficiency and larger effective context windows.
- Code Execution Sandbox: A secure, isolated environment where agents can run generated code without risking the host system.
- Agent Autonomy: The degree to which an agent can make independent decisions and execute complex actions.
- Deterministic Workflow: A workflow that produces the same output every time for the same input.
- Embeddings: Numerical representations of text or data, used here to convert tool definitions into agent knowledge.
- IPython Interpreter Tool: A tool used within the speaker's framework to execute Python code generated by the agent.
Anthropic's Advanced Tool Use: A Paradigm Shift for Agent Efficiency
The video discusses Anthropic's new "Advanced Tool Use" features for Claude, which address the inefficiencies of classic Multi-tool Code Protocol (MCP) servers. Previously, agents would consume 50-100k tokens just on tool definitions before processing any user request. The core idea explored in a previous video, and now adopted by Anthropic, is to let agents generate and run code for tools they need, only exposing the final results. This approach significantly reduces token usage and improves agent performance.
The Problem with Classic MCP and Anthropic's Solution
The traditional MCP approach involves connecting numerous MCP servers and loading dozens of tools per server, leading to substantial token burn (50-100k tokens) before any actual work begins. Anthropic's new features directly tackle this problem, advocating for a mindset shift: "agents should discover and load tools on demand, keeping only what's relevant for the current task."
Key Features of Anthropic's Advanced Tool Use
-
Tool Search Tool:
- Mechanism: Instead of loading all tool definitions upfront, a single "tool search tool" (approximately 500 tokens) is provided. When the agent needs a capability (e.g., "GitHub pull requests"), it calls this search tool with a query. The search tool then returns a small, relevant subset of tool definitions, which are the only ones that enter the agent's context.
- Benefits:
- Token Savings: Quantified by Anthropic, a traditional approach consuming 77k tokens before work begins can be reduced to 8.7k tokens with the Tool Search Tool, preserving 95% of the context window.
- Improved Accuracy: Internal evaluations show accuracy on MCP tasks jumps because Claude is not overwhelmed by a large list of similar tools.
- Implementation: Builders add
tool_search_tool_regaxto their tools array and setdefer_loading: truefor other tools. For MCP servers, the entire server can be deferred by default, with specific "hot tools" (e.g., search files) explicitly loaded. - Analogy: This feature is described as "MCP plus a built-in router," formalizing a manual pattern many builders were already hacking together.
- Recommendation: Essential for setups with multiple MCPs and 50+ tools; less critical for fewer than 10 tools.
- Trade-off: Adds a search step, potentially increasing latency, but offers significant context savings and better tool selection.
-
Programmatic Tool Calling:
- Problem Addressed: Even with optimized tool definitions, intermediate tool results can still pollute the context window.
- Solution: The model writes a Python script to orchestrate tool calls within a code execution sandbox. It processes data inside the script, and only the final, processed JSON output is presented to Claude, rather than all intermediate line items.
- Example: For a task like "Which team members went over their Q3 travel budget?" (using tools like
Get team members,Get expenses,Get budget by level):- Traditional: Multiple sequential natural language tool calls and model round trips to combine data.
- Programmatic: Claude writes a Python script that calls these tools, iterates through results, performs calculations, and prints only the final summary.
- Improvements (Internal Tests):
- Token Savings: Average token usage dropped by 37% on complex research tasks.
- Reduced Latency: Fewer model round trips as one script handles many tool calls.
- Improved Accuracy: Explicit loops and conditionals in code are more reliable than natural language instructions for repetitive tasks.
- Significance: This confirms that "just let your agents run the code" is now a recommended, first-class approach with an API contract, making it more reliable and efficient than manual prompting.
- Implementation: Add
code_executionto the agent's tools. Claude writes orchestration code, which is then executed, and only the final result is provided back to Claude. - Use Cases: Best for agents processing large datasets or managing multi-step workflows with numerous tool calls that require combination or iteration.
-
Tool Use Examples:
- Mechanism: Small, realistic examples of tool calls are directly attached to the tool definition.
- Benefit: The model sees practical examples of how to use a tool, not just its JSON schema, which greatly improves performance, especially for tools with many optional fields or specific conventions.
- Note: While not entirely groundbreaking (examples could be provided in parameter descriptions or system prompts before), this standardizes the approach within the tool definition.
Combined Impact and Challenges
These features collectively enable agents to load only necessary tools, orchestrate them efficiently in code, and use them correctly based on examples. This formalizes and provides API support for patterns previously implemented manually.
However, this increased power comes with significant responsibilities and challenges:
- Less Deterministic: Workflows become less predictable due to potential issues like tool search queries being phrased differently or agents writing incorrect code.
- Harder to Debug: Tracing platforms primarily show the final code output, obscuring intermediate tool calls and data, making it difficult to identify issues within the script.
- Added Latency: The introduction of search and code execution steps can add latency, which might be acceptable for complex workflows but could slow down simple, single-tool tasks.
- Harder to Guarantee Safety: Increased agent autonomy (writing loops, conditionals, combining tools) raises safety concerns. Without proper safeguards, destructive tools could be misused, potentially erasing data or entire tool ecosystems.
- Harder to Implement Correctly: Requires setting up custom, secure sandboxes for code execution and result storage, which is non-trivial infrastructure work.
Real-World Application with Agency AI and Cursor
The speaker demonstrates how these principles are applied in their own agent framework, Agency AI, using the Cursor prompting tool, even with model providers that don't natively support Anthropic's new APIs yet.
Case Study: Cold Email OS Agency
- Problem: A "cold email OS agency" agent, managing campaigns, uses 36 tools and often performs 10-20 tool calls per task. A specific task, "find all inboxes with incorrect settings," required over 20 tool calls and consumed nearly 40,000 tokens. The agent iterated through accounts, checking each individually, instead of processing them in a loop and only logging exceptions.
- Solution using Agency AI/Cursor:
- A new command,
MCP code execution, was added to describe the new pattern to Cursor. - New methods were integrated into the framework to facilitate this pattern.
- Cursor converted the existing MCP server into individual tools.
- A
generate_schema_filefunction was created to save tool schemas into thecampaign_manager_filesdirectory. This directory acts as the agent's knowledge base, with files automatically converted into embeddings. - Cursor tested tools, added code execution tools from the framework, and adjusted agent instructions.
- The agency was deployed on Agency AI.
- A new command,
- Results:
- The agent first searched its knowledge (the embedded tool definitions).
- It then used an
IPython interpreter toolto run functions directly in code, filtering inboxes efficiently. - The final token consumption for the same task dropped significantly from almost 40,000 tokens to 14,000 tokens.
- Conclusion: Agents have become highly proficient at writing code, making it more efficient to let them generate and execute tool orchestration rather than manually building complex tools.
Current Challenges and Platform Solutions:
- Prompting Overhead: Explaining this pattern to an agent still requires significant prompting.
- Infrastructure Overhead: Setting up secure sandboxes for code execution is crucial due to security risks.
- Agency AI Solution: The platform provides secure sandboxes and persistent storage, allowing agents to evolve their skills by saving and reusing files across different chats.
Synthesis and Conclusion
Anthropic's "Advanced Tool Use" features (Tool Search, Programmatic Tool Calling, and Tool Use Examples) represent a significant advancement in agent efficiency and capability. By formalizing the approach of dynamic tool discovery and code-based orchestration, they offer substantial token savings, reduced latency, and improved accuracy. While these features introduce complexities related to determinism, debugging, latency, safety, and infrastructure, they empower builders to create more sophisticated and efficient agents. The speaker's demonstration with Agency AI highlights that similar patterns can be implemented with other models, showcasing the broad applicability and benefits of letting agents leverage their code-writing abilities for tool interaction. The expectation is that more model providers will follow Anthropic's lead in integrating these capabilities directly into their APIs.
AI summaries can miss context or contain errors. Check important details against the original video.