Patrick Dougherty: How to Build AI Agents that Actually Work

AI EngineerAbout 5 min readFeb 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

AI Agents, Autonomous Reasoning, Tool Calls, Retrieval vs. RAG, Reasoning Models, Agent Computer Interface (ACI), Model Selection, Fine-tuning, Abstraction Libraries, Multi-Agent Systems, Manager Agent, Worker Agents, Incentivization.

AI Agent Definition and Key Requirements

Patrick, the former CTO of Rosco, defines an AI agent based on three criteria:

  1. Direction Taking: The agent must be able to receive and act upon directions, whether from humans or other AI, focused on a specific objective.
  2. Tool Access: The agent needs access to at least one tool and be able to receive a response from it.
  3. Autonomous Reasoning: The agent must autonomously decide how and when to use its tools to achieve the objective, not follow a predefined sequence. This distinguishes it from simple prompt chaining.

Focusing on Reasoning Over Knowledge

A major lesson was the importance of enabling agents to think rather than relying solely on pre-existing knowledge. This involves:

  • Discrete Tool Calls for Retrieval: Instead of inserting content into the system prompt (RAG), use specific tool calls to retrieve relevant context during the reasoning process.
  • Example: SQL Query Generation: Providing an agent with all table schemas and columns can overwhelm it, leading to incorrect reasoning and poor query generation. Instead, use tool calls like "search tables," "get table detail," or "profile a column" to guide the agent iteratively.
  • Reasoning Models: Reasoning models can determine if the data needed to answer a question exists. If not, they should be able to report that they cannot find the data, allowing for alternative actions. GPT-4 often attempts to answer questions regardless of data availability.
  • GPT-4 vs. 01 Example:
    • GPT-4: Given a Salesforce schema (accounts, contacts, opportunities) and the question "Write a query to see how many of my customers churned in the last month," GPT-4 attempts to write a query, making assumptions and providing a potentially incorrect SQL statement. It doesn't consider whether the schema contains the necessary information.
    • 01: With the same prompt, 01 reasons through the question and accurately concludes that the provided schema lacks the data to calculate churn.

The Agent Computer Interface (ACI)

The ACI refers to the syntax and structure of tool calls, including the input to the tool and the format of the response. Small tweaks to the ACI can significantly impact agent accuracy and performance.

  • Response Format Matters:
    • GPT-4: Initially, search result payloads were formatted as markdown. The agent would sometimes fail to recognize existing columns in the response. Switching to JSON format solved this issue.
    • Claude: XML format for responses proved more effective than JSON.
  • Key Takeaway: The optimal ACI depends on the specific model being used.

Model Selection and Hallucinations

  • Model as the Brain: The model performs the "thinking" for the agent. A poor model leads to logical fallacies and user dissatisfaction.
  • Prioritize Reasoning in the Core Model: Even if cheaper models are used for sub-tasks, the model responsible for decision-making (which tool to call next) should be highly intelligent. Claude 3.5 is cited as a good balance of speed, cost, and decision-making ability.
  • Learning from Hallucinations: Observe how a model hallucinates to understand its expectations for tool call formats. If a model consistently ignores the JSON schema and provides arguments in a different format, adapt the tool call to match the model's expected format.

Fine-tuning Models

Fine-tuning models for agents was found to be ineffective and even detrimental to reasoning.

  • Reasoning Over Knowledge: If the focus is on reasoning, fine-tuning doesn't improve it.
  • Decreased Reasoning: Fine-tuning can overfit the model to specific task sequences, hindering its ability to reason and make appropriate decisions.

Abstraction Libraries (LangGraph, CrewAI)

The speaker's team chose not to use abstraction libraries for two main reasons:

  1. Availability: When they started, these libraries were not yet publicly available.
  2. Production Considerations: Frameworks like LangGraph presented challenges for production deployment, particularly regarding security and credential management.
    • User Permissions: The agent needed to operate with the end-user's granular permissions within systems like Snowflake, requiring OAuth integration. This was difficult to achieve with existing frameworks.
  • Recommendation: Consider the end goal (production vs. prototype) before becoming too dependent on third-party libraries. Building an agent from scratch doesn't require excessive code.

The Agent's Ecosystem, Not the Prompt, is the Moat

The system prompt is not the most valuable aspect of an agent. The ecosystem around the agent is more important, including:

  • User Experience: How the user interacts with the agent.
  • Connections and Security: The agent's connections to other systems and the security protocols it follows.
  • Key Takeaway: These elements are more time-consuming to build and represent a stronger competitive advantage.

Multi-Agent Systems

Lessons learned from designing multi-agent systems:

  1. Manager Agent: Implement a manager agent to own the final outcome and delegate subtasks to worker agents. This provides more context and specific tool calls to the worker agents.
  2. Team Size: Limit the number of agents working together to between five and eight (similar to the "two-pizza rule"). Larger teams can lead to infinite loops and decreased likelihood of achieving the desired outcome.
  3. Incentivization: Incentivize the manager agent to accomplish the overall objective and manage the worker agents. Avoid forcing worker agents through a discrete set of steps.

Synthesis/Conclusion

Building effective AI agents requires a focus on enabling reasoning through discrete tool calls and careful attention to the Agent Computer Interface. Model selection is crucial, and fine-tuning may not be beneficial. Production considerations, especially security and user permissions, should guide the choice of whether to use abstraction libraries. The ecosystem around the agent, including user experience and secure connections, is more important than the system prompt. In multi-agent systems, a manager agent and incentivization are key to success.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.