OpenAI's AgentBuilder Changes Everything!

By aiwithbrandon

Share:

Key Concepts

  • Agent Kit: OpenAI's new comprehensive tool for building, deploying, and optimizing real-world AI agents.
  • Agent Builder: A user-friendly UI for visually constructing AI agent workflows, built on top of the Agent SDK.
  • Chatkit: A toolkit for integrating agent workflows into applications, enabling custom UIs and interactive experiences.
  • Agent SDK: The underlying framework (Python code) that powers Agent Builder, allowing programmatic agent creation.
  • MCP Tools (Multi-Cloud Platform Tools): Tools that allow agents to connect to external servers and services (e.g., Context 7).
  • RAG (Retrieval Augmented Generation): A technique where an AI model retrieves information from a knowledge base (vector store) to inform its responses.
  • Vector Store: A database that stores embeddings of documents, enabling semantic search for RAG.
  • Guardrails: Mechanisms to ensure AI agents operate safely and effectively, preventing issues like PII leakage, jailbreaking, or hallucinations.
  • Hallucination Guardrail: A specific guardrail that verifies agent claims against trusted documents in a vector store to prevent fabricated information.
  • JSON Schema: A standard for describing the structure and data types of JSON data, used here for defining agent output formats.
  • If/Else Workflows: Conditional logic within Agent Builder to route queries to different agents based on specific criteria.
  • Workflow ID: A unique identifier for a published agent workflow, used to integrate it into applications via Chatkit.
  • Logs & Traces: Monitoring tools within Agent Kit to inspect agent execution, calls, inputs, outputs, and reasoning steps for debugging and optimization.
  • Widgets: Custom UI elements that agents can generate as output, providing rich, interactive responses within applications.
  • Next.js: A popular React framework used in the example for integrating the AI agent application.
  • Context 7: An example MCP tool used for accessing up-to-date code-related documentation.
  • Vercel AI SDK: A tool mentioned for building AI chat applications within Next.js.
  • Shipkit: A platform/course mentioned by the speaker for building real-world AI applications.

OpenAI's Agent Kit: Revolutionizing AI Agent Development

OpenAI's recent Dev Day introduced a suite of new features, with Agent Kit highlighted as the most significant announcement. Agent Kit is designed to simplify the building, deployment, and optimization of production-grade AI agents, making the process feel "like cheating" due to its ease of use. It combines two primary tools: Agent Builder for visual workflow creation and Chatkit for application integration and UI development.


Core Components of Agent Kit

1. Agent Builder: Visual Workflow Creation

Agent Builder provides an intuitive drag-and-drop UI for constructing complex agent workflows. It operates on top of the Agent SDK, which is essentially Python code, abstracting away the complexity for developers. Key functionalities within Agent Builder include:

  • Adding Agents: Core AI entities that perform tasks based on instructions.
  • Adding Tools: Extending agent capabilities with various built-in and custom tools:
    • File Search: For RAG (Retrieval Augmented Generation) against a vector store.
    • Guardrails: To ensure safety and effectiveness (e.g., preventing PII, jailbreaking, hallucinations).
    • MCP Tools: Connecting to external Multi-Cloud Platform servers (e.g., Context 7).
    • Web Search: For general internet queries.
    • Code Interpreter: For executing code.
  • Adding Workflows (If/Else Statements): Implementing conditional logic to route queries to different agents or paths based on specific conditions.
  • Customization Options:
    • Model Selection: Choosing different AI models for reasoning (e.g., GPT-5, with options for "long time" or "little time" thinking).
    • Output Format: Defining how the agent's response is structured (text, JSON, or Chatkit widgets).

2. Chatkit: Application Integration and UI

Chatkit is a toolkit that facilitates the integration of Agent Builder workflows into real-world applications. It allows developers to:

  • Embed agent workflows into applications (e.g., a Next.js app).
  • Build custom UIs for interacting with agents, including interactive widgets.

Example 1: Building a Basic Q&A Workflow with Routing

The first demonstration involved creating a Q&A workflow from scratch, starting simple and adding complexity.

Step-by-Step Process:

  1. Create a New Workflow: Start with a blank canvas in Agent Builder.
  2. Define Core Controls: Understand how to add agents, tools, and workflows.
  3. Code Q&A Agent:
    • Name: "Code Q&A Agent."
    • Instructions: "Your job is to answer questions provided by the user. Every time you answer a code question, use the Context 7 tool to get necessary information."
    • Add Tool (MCP Server):
      • Select "MCP Tool" and "Custom Server."
      • Provide the URL for Context 7 (a tool for accessing up-to-date code documentation).
      • Name the tool "Context 7."
      • Configure tool call approval (e.g., "never" require approval for search).
    • Input Context: Pass the "start" message (initial user question) as input.
    • Preview and Test:
      • Query: "What is the latest version of Vercel AI SDK?"
      • Result: The agent used Context 7, performed reasoning steps, and correctly identified Vercel AI SDK version 5. This demonstrated the agent's ability to use an external tool to fetch specific, up-to-date information.
  4. General Q&A Agent:
    • Name: "General Q&A Agent."
    • Instructions: "Your job is to answer general questions provided by the users and make sure to use the web search tool."
    • Add Tool: Select "Web Search."
    • Input Context: Pass the user question and optionally the entire chat history.
  5. Classifier Agent for Routing:
    • Name: "Classifier."
    • Instructions: "Your job to classify queries as either code or general."
    • Output Format: Change from "text" to "JSON."
    • JSON Schema Generation: Use the built-in "Generate output schema" feature to define the output as an enum with categories "code" or "general." This simplifies complex JSON schema creation.
    • Input Context: Pass the user query.
  6. If/Else Workflow for Conditional Routing:
    • Add an "If/Else" node.
    • Condition 1 (Code): If category (from the Classifier agent's output) equals "code," route to the "Code Q&A Agent."
    • Condition 2 (Else): Otherwise, route to the "General Q&A Agent."
    • Preview and Test Routing:
      • Query 1: "Can you please tell me what is the latest version of NextJS?"
        • Result: Classifier identified as "code," routed to Code Q&A, which then answered the question.
      • Query 2: "What is the weather today in Atlanta, Georgia?"
        • Result: Classifier identified as "general," routed to General Q&A, which used web search to provide the weather.

This example showcased the power of combining multiple agents, external tools, and conditional logic to create a sophisticated, routed Q&A system in minutes.


Example 2: Building a RAG Agent with Guardrails and Chatkit Integration

The second example focused on building a more complex RAG agent and integrating it into a Next.js application.

Step-by-Step Process:

  1. Create a New Workflow: Start a fresh workflow.
  2. RAG Query Agent:
    • Name: "Rag Query Agent."
    • Instructions: "Your job is to be a query agent where you are going to answer questions that students might have had about the coaching calls during the weekly Shipkit coaching calls. Search up the information for their question and provide a well-informed answer." (The vector store contains transcripts of coaching calls).
    • Add Tool (File Search):
      • Select "File Search."
      • Select Vector Store:
        • Navigate to OpenAI's platform to create a new vector store (e.g., "Brandon Hancock").
        • Upload markdown files containing coaching call transcripts (e.g., call1.md, call2.md).
        • Copy the Vector Store ID.
        • Paste the ID into Agent Builder.
    • Input Context: Pass the user query.
    • Preview and Test:
      • Query: "Can you please tell me when the last coaching call was for shipkit.ai?"
      • Result: The RAG agent performed queries against the vector store, retrieved relevant information, and correctly stated the last call was on October 7th, including a link to the recording. This demonstrated real-world RAG capabilities.
  3. Implementing Guardrails (Hallucination Prevention):
    • Add Guardrail Node: Drag a "Guardrail" node onto the canvas.
    • Configure Hallucination Guardrail:
      • Select "Hallucinations."
      • Specify the same vector store used by the RAG agent.
      • Choose a model and confidence level (default is fine for quick verification).
    • Route Agent Output to Guardrail: Connect the RAG agent's output to the Guardrail.
    • Conditional Routing (Guardrail Pass/Fail):
      • If Pass: Route to a "Result Agent" with instructions: "Please return the query answer to the user."
        • Context: Pass "safe text" (verified answer from guardrail) and "initial question."
      • If Fail: Route to an "Unsafe Answer Agent" with instructions: "The user's initial query produced hallucinations. Here is the initial query. Here is the reason why the answer is not valid. Here is the hallucination type as well."
        • Context: Pass "initial query," "reasoning" (from guardrail), and "hallucination type" (from guardrail).
    • Preview and Test Guardrail:
      • Query (designed to fail): "What time did Brandon walk his dog yesterday?"
      • Result: The RAG agent searched, found no relevant information, the Guardrail failed (as expected), and the "Unsafe Answer Agent" provided a detailed explanation: "The user is asking about walking the dog but the documents never mention any information about his dogs. Therefore, the factual claim or any questions is going to be unsupported." This demonstrated effective hallucination prevention.

Integrating Agents with Next.js using Chatkit

The final part of the demonstration focused on deploying the RAG agent into a Next.js application using Chatkit.

Step-by-Step Process:

  1. Chatkit Starter Template: Use the provided OpenAI Chatkit starter app, which includes necessary packages for Chatkit integration.
  2. Publish Agent Workflow:
    • In Agent Builder, click "Publish" in the top right.
    • Name the workflow (e.g., "Rag Coaching").
    • After publishing, click "Code" to access the Workflow ID. This ID is crucial for connecting the application to the agent.
    • Note: The "Code" view also reveals the underlying Python code generated by Agent Builder, confirming that Agent Builder is a UI on top of the Agent SDK.
  3. Integrate into Next.js Application:
    • Copy the Workflow ID from Chatkit.
    • In the Next.js starter app, update the public workflow ID variable.
    • Provide an OpenAI API key.
    • Run npm install and npm run dev.
  4. Live Application Test:
    • Access localhost:3000.
    • Query: "What time was Brandon's latest coaching call?"
    • Result: The Next.js application successfully interacted with the deployed RAG agent, performing the same query and returning the correct answer, demonstrating seamless production deployment.

Monitoring and Debugging:

  • Logs: The Agent Kit platform provides "Logs" to view Chatkit threads, showing user queries and agent responses.
  • Traces: "Traces" offer a deep dive into workflow execution, detailing agent calls, configurations, system instructions, inputs, and outputs at each step, which is invaluable for debugging and understanding agent behavior. This feature is highlighted as a significant advantage, as such tracing capabilities often require extensive setup or third-party tools in other frameworks.

Advanced Chatkit Features: Custom Widgets

Agent Kit extends beyond plain text responses by allowing agents to generate rich, interactive UI elements called widgets.

Step-by-Step Process:

  1. Change Output Format: In the "Result Agent" (or any agent), change the output format from "text" to "Chatkit Widget."
  2. Create Widget:
    • Click "Add Widget" and then "Create a new widget."
    • Instructions: Provide natural language instructions for the desired UI (e.g., "Create a really clean, simplistic UI element that will contain the search results from a rag agent. All we need to have is a nice header and a body. Make it clean, minimal for our users.").
    • AI Generation: The AI generates the widget definition, specifying required inputs (e.g., header, body).
    • Download/Upload: Download the generated widget definition and upload it back to Agent Builder.
  3. Pass Data to Widget: Connect the agent's output (e.g., the safe text from the guardrail) to the widget's inputs (e.g., map the answer to the body and a title to the header).
  4. Preview and Test Widget:
    • Query: "Can you please tell me what is the chat SAS application?"
    • Result: After RAG and guardrail processing, the agent returned the answer formatted within the custom widget, displaying a header and body.
  5. Next.js Integration: The same widget output seamlessly appeared in the Next.js application, demonstrating the ease of building and deploying custom UIs.
  6. Customization: Widgets can be highly customized with tables, colors, and interactive actions (e.g., copy, send buttons). Pre-built widget examples (weather, to-do calendar) are also available, and their source code can be inspected and modified.

Conclusion

OpenAI's Agent Kit, comprising Agent Builder and Chatkit, represents a significant leap in simplifying the development and deployment of real-world AI agents. It offers a powerful visual interface for building complex workflows, robust tools for RAG and safety (guardrails), and seamless integration with applications through Chatkit, including the ability to generate rich, interactive UI widgets. While acknowledging some initial reliability issues with new technologies, the platform's capabilities for rapid prototyping, deployment, and detailed tracing are highlighted as game-changers, drastically reducing the time and effort traditionally required for such advanced AI applications.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video