THE SUMMARYAI-generated
Key Concepts
- Agents: Models using tools in a loop to accomplish tasks independently.
- Environment: The context in which the agent operates, including available tools.
- System Prompt: Instructions given to the agent defining its goals and behavior.
- Reasonable Heuristics: General principles or guidelines provided to the agent to improve performance.
- Tool Selection: The process of determining which tools an agent should use for specific tasks.
- Interleaved Thinking: The ability of the model to reflect on the results of tool calls and adjust its strategy accordingly.
- Context Window Management: Techniques for extending the effective context window of the model, such as compaction and sub-agents.
- Compaction: A tool that summarizes or compresses the context window to allow the agent to continue running without running out of context.
- Sub-agents: Delegating tasks to other agents to manage context window limitations.
- Evaluations (Evals): Methods for systematically measuring the performance of an agent and identifying areas for improvement.
- LLM as Judge: Using a large language model to evaluate the output of an agent based on a defined rubric.
- Towen: An open-source benchmark for evaluating whether agents reach the correct final state.
What are Agents?
- Anthropic defines agents as "models using tools in a loop."
- Agents are given a task and allowed to work continuously, using tools as needed.
- They update their decisions based on information from tool calls and work independently until the task is complete.
- Key components:
- Environment: Where the agent works.
- Tools: Resources available to the agent.
- System Prompt: Instructions for the agent.
- Simpler system prompts are generally better, allowing the model to perform optimally.
When to Use Agents
- Agents are best suited for complex and valuable tasks.
- Avoid using agents for simple tasks where a step-by-step process is clear.
- Checklist for Agent Use:
- Complexity: Is the task complex and lacking a clear path to completion?
- Value: Is the task high-value and revenue-generating?
- Doability: Can the necessary tools and information be provided to the agent?
- Cost of Errors: Are errors easily detectable and recoverable? If not, consider a human-in-the-loop approach.
Examples of Agent Use Cases
- Coding:
- Transforming a design document into a pull request (PR).
- High-value task that saves engineers time.
- Search:
- Errors are recoverable through citations and double-checking.
- Computer Use:
- Recoverable errors (e.g., re-clicking).
- Data Analysis:
- Extracting insights or creating visualizations from data with unknown formats or errors.
Prompting for Agents: Best Practices
- Think Like Your Agent:
- Develop a mental model of the agent's environment (tools and responses).
- Simulate the agent's process to identify potential confusion or limitations.
- If a human can't understand what the agent should be doing, the AI won't either.
- Give Reasonable Heuristics:
- Prompt engineering is "conceptual engineering."
- Define concepts and behaviors for the model to follow.
- Example: Cloud Code has the concept of irreversibility (avoiding harmful actions).
- Be clear about concepts and consider edge cases.
- Example: Provide budgets for tool calls (e.g., under five for simple queries, up to 15 for complex queries).
- Articulate principles clearly, as you would for a new intern.
- Tool Selection is Key:
- Be clear about which tools the agent should use for different tasks.
- Provide explicit principles for tool usage in specific contexts.
- Example: Default to searching Slack for company-related information.
- Guide the Thinking Process:
- Prompt the agent to plan its search process in advance (complexity, tool calls, sources, success criteria).
- Use interleaved thinking to reflect on the quality of search results and verify information.
- Agents are More Unpredictable:
- Most changes will have unintended side effects due to autonomous operation.
- Example: Telling the agent to keep searching until it finds the perfect source may lead to endless searching.
- Help the Agent Manage Its Context Window:
- Use techniques like compaction to summarize the context window.
- Write to external files to extend memory.
- Use sub-agents to delegate tasks and compress results.
- Let Claude Be Claude:
- Start with a bare-bones prompt and tools, then iterate based on observed issues.
- Claude is often surprisingly capable.
- Tool Design:
- Ensure tools are well-designed with simple, accurate names and descriptions.
- Test tools thoroughly.
- Avoid giving the agent multiple tools with similar names or descriptions.
Demo Example: How Many Bananas Fit in a Rivian R1S?
- The agent uses web search to find the cargo capacity of the Rivian R1S and the dimensions of a banana.
- It reflects on the results and converts measurements.
- It runs calculations to estimate the number of bananas that can fit.
- The model estimates approximately 48,000 bananas, which is roughly correct.
Evaluations (Evals) for Agents
- Evals are crucial for systematically measuring progress.
- Evals are more difficult for agents due to their long-running and unpredictable nature.
- Tips for Easier Evals:
- Start with a small eval and iterate.
- Use realistic tasks.
- Use LLM as judge with a clear rubric.
- Human evals are essential.
- Examples of Evals:
- Answer Accuracy: Use an LLM as judge to determine if the answer is accurate.
- Tool Use Accuracy: Evaluate if the correct tools were used at the right times.
- Towen: Evaluate whether the agent reaches the correct final state (e.g., database updates, file modifications).
Building Prompts for Agents: Iterative Approach
- Start Simple: Begin with a short, basic prompt.
- Test and Observe: Run the agent and observe its behavior.
- Identify Edge Cases: Collect test cases where the model fails or succeeds.
- Iterate and Refine: Add instructions and examples to the prompt to address edge cases and improve performance.
- Monitor Progress: Continuously evaluate the agent's performance and adjust the prompt as needed.
Few-Shot Examples and Chain of Thought
- Traditional prompting techniques like few-shot examples and chain of thought are less effective for state-of-the-art frontier models and agents.
- These techniques can limit the model's flexibility and creativity.
- Instead, focus on guiding the agent's thinking process and providing clear instructions on how to use its thinking.
- Give examples, but avoid being too prescriptive.
Synthesis/Conclusion
The presentation provides a comprehensive guide to prompting for agents, emphasizing the importance of understanding the agent's environment, providing reasonable heuristics, and guiding the thinking process. It highlights the iterative nature of prompt engineering and the need for robust evaluation methods. By following these best practices, developers can create effective and reliable agents for a wide range of complex tasks.
AI summaries can miss context or contain errors. Check important details against the original video.





