Anthropic’s Set 46 Release: Programmatic Tool Calling & Performance Improvements
Key Concepts:
- Programmatic Tool Calling: Invoking tools via code execution within a sandbox environment, rather than relying on traditional JSON-based tool calling.
- MCP (Model Context Protocol): A protocol for connecting LLMs to various tools and data sources.
- Context Window Problem: The limitation of LLMs to process only a finite amount of text, leading to potential information overload and reduced performance.
- Context Engineering: The practice of carefully curating the information within the context window to maximize relevance and efficiency.
- Sandbox Environment: A secure, isolated environment for executing code without affecting the main system.
- Dynamic Filtering: Using code execution to post-process web search results, retaining only the most relevant information before it enters the context window.
- Browser Comp & Deep Search QA: Benchmarks used to evaluate the performance of agents in web search and information retrieval tasks.
- Price-Weighted Tokens: A metric that considers both the number of tokens used and the cost per token, reflecting the overall expense.
1. Introduction to Programmatic Tool Calling
With the release of Set 46, Anthropic has introduced developer tools focused on cost reduction and performance enhancement for agents, primarily through programmatic tool calling. This feature allows agents to invoke specific tools by writing code, rather than relying on the agent attempting to load all tool definitions into the context window. This approach improves both token efficiency and accuracy. The speaker highlights Anthropic’s history of setting industry trends, citing the adoption of MCP (Model Context Protocol) and agent skills by other companies following Anthropic’s initial implementation. The comparison to OpenAI’s chat completion API further emphasizes this pattern of industry-wide adoption.
2. The Context Window Problem & The Need for Programmatic Tool Calling
The core issue driving the need for programmatic tool calling is the context window problem. As agents utilize protocols like MCP, tool definitions, input/output data from tool calls, system prompts, user memory, and user messages all compete for space within the limited context window. This “pollution” of the context window with unnecessary information hinders performance. Context engineering – the practice of providing only useful information – becomes crucial. Traditional tool calling exacerbates this issue, as the results of each tool call are added to the context, creating a chain of information.
Programmatic tool calling addresses this by having the agent write code to invoke tools within a sandbox environment. Only the final summary or answer is passed back to the agent, keeping intermediate steps contained and reducing the overall token count.
3. Timeline of Development & Industry Exploration
The concept of programmatic tool calling isn’t exclusive to Anthropic. The speaker outlines a timeline:
- September 2025: Cloudflare published a report ("Code Mode, the better way to use MCP") demonstrating token savings of 30-80% through a sandboxed approach.
- November 2025: Anthropic published “Code execution with MCP,” reaching similar conclusions as Cloudflare.
- Late November 2025: Anthropic released advanced tools, including tool search (for finding specific tools within an MCP server).
- Post-Release: Open-source implementations emerged, including support in Blocks Goose Agent and LightLLM (with native support across providers).
4. Performance Results & Benchmarks
Anthropic’s Set 46 release includes improved web search and dynamic filtering capabilities, both powered by programmatic tool calling. Previously, models would “dump” all search results into the context window. Now, Cloud can write and execute code to filter these results before they are added to the context, improving accuracy and token efficiency.
Results from two benchmarks are presented:
- Browser Comp: Tests an agent’s ability to navigate multiple websites to find obscure information. Sonnet 46 saw a 13% improvement, while Opus 46 saw a 16% improvement.
- Deep Search QA: Tests the ability to find all correct answers to a question via web search. Sonnet 46 showed an F1 score improvement from 52% to 59%, and Opus saw an almost 8% improvement.
Notably, token cost reduction isn’t guaranteed. While Sonnet 46 saw a decrease in price-weighted tokens on both benchmarks, Opus 46 experienced an increase due to the model writing significantly more code for filtering.
5. Practical Implementation & API Usage
Using the search API with Set 46 requires no changes to existing code. Anthropic automatically leverages the new dynamic filtering capabilities when data fetching is enabled. The speaker highlights the availability of detailed documentation and examples, demonstrating how to define tools with input and output schemas. Instead of function calling, Cloud will now write code to execute these tools.
6. Future Outlook & Industry Standard
The speaker predicts that programmatic tool calling will likely become an industry standard, similar to MCP and agent skills. They encourage viewers to share specific use cases and express interest in detailed tutorials for integrating this feature with other coding agents.
Notable Quote:
“LLMs are trained on billions of lines of code, especially with coding agents. You can see that they can produce and understand code but barely any synthetic JSON tool calling formats. So you want to do or let the agent do what they are good at and that is writing code.” – Speaker, emphasizing the natural fit between LLMs and code execution.
Conclusion:
Anthropic’s Set 46 release introduces a significant advancement in agent development with programmatic tool calling. By leveraging the LLM’s inherent coding capabilities and utilizing a sandboxed execution environment, this feature promises to reduce token costs, improve accuracy, and address the limitations of the context window. While not a guaranteed cost saver in all cases (as demonstrated by Opus’s results), the potential benefits and the speaker’s prediction of industry-wide adoption make this a crucial development to watch for anyone working with LLMs and agents.
AI summaries can miss context or contain errors. Check important details against the original video.