Anthropic Just Changed How Agents Call Tools. I Stole It for My Qwen3.5 Agent

By The AI Automators

Share:

Key Concepts

  • Tool Search (Deferred Loading): A design pattern where tools are not loaded into the context window upfront, but are discovered and loaded dynamically via a registry search.
  • Programmatic Tool Calling (Code Mode): An approach where an LLM writes and executes a script (e.g., Python) to perform multiple tool calls internally, rather than performing them sequentially through the LLM's chat loop.
  • Tool Bridge: A secure architectural component that allows an isolated sandbox to communicate with external APIs/services via a controlled FastAPI backend.
  • MCP (Model Context Protocol): A standard for connecting AI agents to data sources and tools.
  • GVisor: A container runtime sandbox that provides stronger isolation than standard Docker by intercepting system calls.
  • Context Bloat: The phenomenon where excessive tool definitions and intermediate execution results consume the LLM's context window, reducing performance and increasing costs.

1. Advanced Tool Calling Methods

The video explores two primary methods to optimize AI agent performance, originally highlighted by Anthropic but applicable to any framework:

  • Tool Search: Instead of loading all available tools (e.g., 60+ tools from multiple MCP servers) into the context, the agent only loads a "Tool Search" tool. When a task requires a specific function, the agent searches the registry, retrieves the relevant schema, and loads only that specific tool into the context. This significantly reduces token usage (e.g., from 13,000 to 6,300 tokens in the demo).
  • Programmatic Tool Calling: This method shifts the burden of execution from the LLM's chat loop to a script. The LLM generates a script (e.g., a for loop to fetch expenses for 20 team members) that runs in a sandbox. This avoids the "ping-pong" effect of multiple LLM turns, reducing token consumption and preventing intermediate results from polluting the context.

2. Architecture and Security

The system utilizes a custom Python/React application with the following security and execution flow:

  • Isolated Sandbox: Code execution occurs within a Docker container (using LLM-sandbox).
  • Secure Tool Bridge: The sandbox has no direct internet access. It communicates with the host FastAPI app via a session-authenticated bridge. The host app validates the tool schema and executes the actual API calls.
  • GVisor Integration: To mitigate the security risks of shared kernels in standard Docker, the author recommends using GVisor for enhanced isolation.
  • TypeScript/Python Stubs: To improve reliability, MCP schemas are converted into language-specific stubs (e.g., Python functions). This allows the LLM to interact with tools using standard, well-understood syntax rather than complex JSON structures.

3. Comparative Performance

  • Traditional vs. Programmatic: In a budget analysis task, the traditional approach required 56 tool calls and 76,000 tokens, often failing to provide a comprehensive answer. The programmatic approach, while iterative, achieved higher accuracy with significantly fewer LLM-level interactions.
  • Model Comparison: The author tested both Claude Haiku and Qwen 3.5 (27B parameter model). Qwen 3.5 performed efficiently, requiring fewer tokens (45,000) to reach the correct answer compared to Haiku in the specific test case.

4. Key Arguments and Perspectives

  • Strategic Layering: The author argues that "Anthropic hasn't killed tool calling." Instead, developers should layer these features strategically:
    • Use Tool Search to solve context bloat from definitions.
    • Use Sandboxing/Programmatic Calling to solve context bloat from intermediate results.
    • Use Tool Use Examples (multi-shot prompting) to solve parameter formatting errors.
  • Data Processing Responsibility: A central question posed is whether the LLM should perform ad-hoc data processing or if developers should pre-create "Skills" (verified scripts) that the LLM simply triggers. The author suggests that for complex, repetitive tasks, pre-built scripts are more reliable than relying on the LLM to "one-shot" the logic.

5. Notable Quotes

  • "The key aspect of the tool search tool is you don't load everything up front. You defer the loading and allow the agent to search for it."
  • "Should your LLM actually be doing the data processing ad hoc like this, or should it simply just be relaying the information from a pre-created script?"

6. Synthesis and Conclusion

The transition from traditional, sequential tool calling to deferred loading and programmatic execution represents a shift toward more robust, production-grade AI agents. By offloading logic to sandboxed scripts and dynamically managing the tool registry, developers can drastically reduce context bloat and improve the accuracy of complex multi-step tasks. The most effective systems are those that combine these advanced patterns with secure, isolated execution environments, ensuring that the LLM acts as a coordinator rather than a bottleneck for data processing.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video