Building tools for agents — with agents

Prompt EngineeringAbout 5 min readSep 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Agentic Systems
  • Tool Implementation
  • Context Engineering
  • Prompt Engineering for Tools
  • Tool Selection
  • Tool Grouping (Namespacing)
  • Token Efficiency
  • Evals (Evaluation Datasets)
  • Agentic Loop for Tool Improvement

1. Software Engineering Shift with Agentic Systems

  • Traditional Software Engineering: Deterministic functions with defined inputs and outputs.
  • Agentic Systems: Input is a natural language query. The agent decides on steps, including tool calls, internal knowledge, or follow-up questions. The same question can lead to different steps each time.
  • Agentic systems are often integrated into traditional software engineering stacks, requiring a rethinking of implementation.

2. Four-Step Process for Building Good Tools (Brief Overview)

  • Initial Prototype: Implement core functionality.
  • Evals: Run evaluations on the implemented tools.
  • Collaboration with Coding Agent: Use coding agents to improve tool description and implementation.
  • Iterative Evals: Continuously evaluate and improve based on metrics.

3. Key Principles for Tool Implementation

  • Choose the Right Tools: Focus on meaningful outcomes for the agent, considering the context window.
  • Group Tools Together (Namespacing): Organize similar tools to avoid agent confusion.
  • Context Engineering: Ensure meaningful input to the tool and meaningful context returned to the LLM.
  • Token Efficiency: Avoid filling the context window with useless information.
  • Prompt Engineering: Effectively craft tool descriptions to aid LLM selection.
  • Refinement over Quantity: Small improvements can yield significant results.

4. Choosing the Right Tools: Context Window Considerations

  • Common Error: Simply wrapping existing software functions/APIs as tools.
  • Context Window Focus: Design tools that generate meaningful outcomes within the LLM's context window.
  • Address Book Search Example:
    • Bad Approach: List Contacts (returns all 500 contacts, potentially filling the context window with useless information).
    • Good Approach: A tool that searches and returns a small, relevant subset of contacts (like a simple RAG system).
  • Human Interaction Analogy: Design tools that mimic how humans interact with them (search, filter, navigate).
  • More Examples:
    • Instead of List Users, List Events, Create Events, implement Schedule Event (finds availability and schedules).
    • Instead of Read Logs, implement Search Logs (returns relevant logs and surrounding text).

5. Grouping Tools (Namespacing)

  • Problem: Agents can get confused when choosing between many tools with similar functionality across different MCPs (Managed Cloud Providers).
  • Solution: Group tools by adding prefixes (or suffixes) based on the MCP or service they belong to (e.g., Slack Messages: Search, Jira: Create).
  • Prefix vs. Suffix: Experiment with both prefixes and suffixes to see which yields better tool usage performance.

6. Context Engineering: Tool Response Optimization

  • High Signal Information: Return only meaningful information to the agent.
  • Descriptive Fields: Use fields like name, image URL instead of low-level technical identifiers (UUIDs, encrypted URLs).
  • Structured Outputs: Return structured formats (XML, JSON, Markdown) that preserve both low-level and high-level details.
  • Model-Specific Formats: Choose the data format (XML, JSON, Markdown) based on the model's training (e.g., Claude likes XML).
  • Token Optimization:
    • Use pagination, range selection, filtering, or curation mechanisms for longer contexts.
    • Encourage agents to pursue token-efficient strategies (small, targeted searches).
    • Claude restricts tool responses to 25,000 tokens by default.

7. Prompt Engineering for Tool Description

  • Importance: Often overlooked but crucial for tool selection.
  • Focus: The agent primarily relies on the tool's description and input/output specifications.
  • Intern Analogy: Provide the type of information a new intern would need to understand the tool, assuming no prior knowledge.
  • Key Elements:
    • Specialized query formats.
    • Definitions of niche terminologies.
    • Relationships between underlying resources.
    • Explicit instructions to avoid ambiguity.
  • Communication is Key: Clear communication is essential for successful agentic systems.

8. Improving Tools with Agents: Evals and Agentic Loops

  • Real-World Evals: Use real-world test cases, not synthetic examples.
  • Comprehensive Information: Provide all necessary documentation to the agent.
  • Complex Scenarios: Test tools in combination, not in isolation.
  • Weak vs. Strong Task Example:
    • Weak: "Schedule a meeting with Jane next week." (single tool selection)
    • Strong: "Schedule a meeting with Jane next week to discuss our latest project. Attach the notes from our latest project planning meeting and reserve a conference room." (multiple tool selection and planning)
  • Agentic Loop: Run evals in a loop, gather feedback, and use agents to improve the tools.
  • Performance Improvement: Using agents to improve tools can lead to substantially better performance.

9. Browserbase and Stagehand

  • Browserbase: A cloud platform for hosting browser automation tools like Puppeteer and Playwright.
  • Stagehand: Browserbase's open-source AI browser automation framework.
    • Enables browser automation by defining actions and assigning tasks to agents.
    • Comes with a powerful SDK, pre-existing agents (e.g., computer use agent), and Playwright compatibility.
    • Allows interaction with Playwright through natural language.
    • Enables computer use agents.
    • Allows scraping data from websites using custom schemas.

10. Conclusion

Building effective tools for agents requires a shift in thinking from traditional software engineering. Key principles include choosing the right tools for the task, grouping tools logically, optimizing context, crafting effective prompts, and continuously evaluating and improving tools using agentic loops. Focusing on real-world use cases and avoiding ambiguity in tool descriptions are crucial for success. Browserbase and Stagehand offer valuable tools for AI agents interacting with web browsers.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.