THE SUMMARYAI-generated
Key Concepts
- Agents: Autonomous systems that accomplish tasks on behalf of users.
- Workflows: A sequence of steps executed to meet a user goal.
- LLM (Large Language Model): The core of an agent, responsible for reasoning and decision-making.
- Tools: Capabilities expanded by function calls, allowing agents to interact with external systems.
- Instructions: System instructions, guidelines, and guardrails that control agent behavior.
- Agent SDK: OpenAI's framework for building agents.
- Evals: Evaluation datasets used to measure and improve agent performance.
- Guardrails: Mechanisms to control data privacy risks or reputational risks.
- Single-Agent Systems: An agent with a set of instructions, tools, guardrails and hooks.
- Multi-Agent Systems: Multiple agents controlling different parts of workflows.
- Manager Pattern: A central agent orchestrates other agents as tools.
- Decentralized Pattern: Agents work autonomously with handoffs of control.
What is an Agent?
- OpenAI defines agents as systems that independently accomplish tasks.
- Key characteristics:
- Autonomous decision-making capabilities.
- Access to tools for interacting with external systems.
- Workflows are sequences of steps to achieve a user goal.
- Not every LLM-based application requires an agent (e.g., simple chatbots, sentiment classifiers).
When to Build an Agent
- Not every LLM-based solution needs an agent.
- Criteria for needing an agentic solution:
- Complex Decision Making: Nuanced reasoning is required.
- Difficult-to-Maintain Rules: Rule-based systems become too complex.
- Heavy Reliance on Unstructured Data: Natural language, document extraction.
- Validate that the use case meets these criteria before committing to an agent.
Components of an Agentic System
- Model (LLM): Powers the agent's reasoning and decision-making.
- Tools: Capabilities expanded by function calls.
- Instructions: System instructions, guidelines, and guardrails.
Model Selection
- Optimize for performance versus cost and latency.
- Recommendation:
- Build a prototype with the most capable model to establish a performance baseline.
- Replace it with weaker models to assess the impact.
- Focus on meeting accuracy targets while optimizing for cost and latency.
Tools
- Tools are categorized into:
- Data: Retrieve context and information (e.g., RAG pipelines).
- Action: Interact with systems (e.g., send emails, update CRM records).
- Orchestration: Agents serving as tools to other agents.
- Implementation:
- Tools are Python functions with a
@tooldecorator. - Detailed docstrings describe the tool's function, inputs, and outputs.
- Limit the number of tools per agent; split into categories if necessary.
- Tools are Python functions with a
Instructions
- System messages that control the behavior of the model.
- Best practices:
- Use existing documents (operating procedures, workflows, policies).
- Prompt agents to break down tasks into subtasks.
- Provide clear instructions; avoid assumptions.
- Capture edge cases (incomplete user information).
- Iteratively refine instructions, tools, and models based on failure cases.
Orchestration Patterns
- Single-Agent Systems: A single model executes workflows in a loop.
- User input -> Reasoning loop (tool selection, execution) -> Output.
- Implemented using
runner.runin the Agent SDK.
- Multi-Agent Systems: Multiple agents coordinate workflow execution.
- Guidelines for splitting into multi-agent systems:
- Complex decision logic (many conditional statements).
- Too many tools (ambiguous names, overlapping functions).
- Limit tools to around 10 per agent.
- Guidelines for splitting into multi-agent systems:
- Manager Agent (Orchestrator): A central agent controls other agents as tools.
- All communication goes through the orchestrator.
- Decentralized Agent: Agents work autonomously with handoffs of control.
- Triage agent directs the user to a specialized agent.
- Example: Technical support agent, sales assistant agent, order management agent.
- Both patterns can be modeled as graphs.
Guardrails
- Critical for customer-facing applications to manage data privacy, risks, or reputational risks.
- Implemented independently of the agentic system.
- Types of guardrails:
- Relevancy classifier
- Safety classifier
- PPI filter
- Moderation guardrail
- Tool safeguard (SQL injections, prompt injections)
- Output validation
- Iterative refinement: Observe system behavior and update guardrails.
- Implementation:
- Input and output guardrails triggered by specific tripwires.
- Implemented as Python functions with decorators.
Additional Recommendations
- Evals: Robust evaluation datasets are essential. Start small and expand based on system behavior.
- Metrics: Track application-specific metrics (accuracy, recall).
Conclusion
The OpenAI practical guide provides a comprehensive overview of building agents, emphasizing the importance of clear instructions, appropriate tool selection, and robust guardrails. The guide advocates for iterative refinement and a focus on practical considerations like cost and latency. The convergence of ideas from different labs suggests a common path forward in the development of agentic systems.
AI summaries can miss context or contain errors. Check important details against the original video.





