How AI Uses Tools

By Don Woodlock

Share:

Key Concepts

  • Agentic AI: AI systems characterized by multi-step workflows, tool calling, and non-deterministic decision-making.
  • Tool Calling: The ability of an LLM to interact with external APIs, databases, or software to perform actions beyond its internal training data.
  • Non-deterministic Workflow: A process where the LLM dynamically decides the next step based on environmental feedback rather than following a rigid, pre-programmed sequence.
  • XML Prompting: The original method for tool calling where LLMs were instructed to output structured XML to request external data.

1. Characteristics of Agentic Workflows

The speaker defines an agentic workflow through three core pillars:

  • Multiple LLM Calls: Workflows involve multi-step processes, often utilizing different agents (e.g., a writer, a critic, and a finalizer).
  • Tool Calling: The model acts as an interface to external systems (APIs, browsers, databases).
  • Non-deterministic Decision Making: Unlike traditional software that follows a fixed "if-then" logic, the LLM determines the sequence of actions based on the results of previous tool calls.

2. The Evolution of Tool Calling

The speaker outlines four distinct phases in the development of tool calling:

  • Phase 1: The XML Era (The "Web Search" Origin): To overcome the "training date problem" (where models lacked current information), developers instructed LLMs to output XML when they needed external data. The application would parse the XML, execute a web search API call, and feed the results back to the LLM.
  • Phase 2: Framework Integration: Tools like LangChain emerged to abstract the "messiness" of manual XML parsing. These frameworks handled the formatting, parsing, and execution of tool calls automatically.
  • Phase 3: Native API Support: Companies like OpenAI and Anthropic integrated tool calling directly into their APIs. Models are now specifically trained to recognize and execute tool calls, removing the need for manual XML prompting.
  • Phase 4: User-Centric Tool Selection (Future): The upcoming phase involves systems where users can select from a "grab bag" of available tools, and the AI autonomously manages the execution to complete the task.

3. Step-by-Step Workflow Process

The speaker describes the operational loop of a tool-calling agent:

  1. Input: The application sends the user’s query and a list of available tools to the LLM.
  2. Decision: The LLM evaluates if it can answer directly or if it needs external data.
  3. Request: If external data is needed, the LLM outputs a request (originally XML, now via API structure) specifying the tool and arguments.
  4. Execution: The application parses the request, executes the API call (e.g., a Google search), and retrieves the data.
  5. Feedback Loop: The application sends the search results back to the LLM.
  6. Resolution: The LLM processes the new information to either perform another tool call or provide a final answer to the user.

4. Key Arguments and Perspectives

  • Overcoming Static Knowledge: The primary argument for tool calling is that it effectively bypasses the limitations of an LLM’s training cutoff date, allowing for real-time, accurate information retrieval.
  • Programmer vs. User Orientation: The speaker notes that current tool-calling implementations are highly "programmer-oriented," requiring developers to build the tools themselves. The shift toward "user-oriented" systems will be the next major milestone in AI accessibility.

5. Synthesis and Conclusion

The invention of tool calling was a pivotal moment in AI development, transforming LLMs from static text generators into active agents capable of interacting with the real world. By moving from manual XML parsing to native API integration, the industry has enabled more complex, non-deterministic workflows. This evolution is the foundation of the current "agentic AI explosion," shifting the paradigm from simple question-answering to task-oriented automation.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video