Key Concepts
Responses API, Chat Completions API, File Search, RAG (Retrieval Augmented Generation), Computer Use, Code Interpreter, Custom Tools, Web Search Tool, Streaming, OpenAI Platform, Gradio, LM Studio, Agentic AI Systems.
Responses API: A Powerful Tool Replacing Chat Completions
OpenAI's Responses API is presented as a significant upgrade to the Chat Completions API, designed to power a wider range of AI applications. The key advantage of the Responses API is its support for features not available in the older Chat Completions API, including:
- File Search: Enables the AI to search through uploaded files (e.g., PDFs) to answer questions, facilitating Retrieval Augmented Generation (RAG).
- Computer Use: Allows the AI to control the user's computer, automating repetitive tasks.
- Code Interpreter: (Coming Soon) Will enable the AI to generate and execute code.
The speaker emphasizes that understanding the Responses API is crucial for anyone building AI applications.
Basic Chatbot Creation
The video provides a step-by-step guide to creating a basic chatbot using the Responses API.
1. Installation:
- Install necessary Python packages using
pip:pip install openai geopy gradioopenai: The core OpenAI library.geopy: Used for geolocation in custom tools.gradio: Used for creating a user interface.
2. API Key Setup:
- Obtain an OpenAI API key from platform.openai.com.
- Set the API key in your code.
3. Code Implementation (app.py):
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY")
response = client.responses.create(
model="gpt-4-turbo-preview",
input="Write a one sentence bedtime story about a unicorn."
)
print(response.output.text)
- The code imports the
OpenAIlibrary. - It initializes an OpenAI client with your API key.
- It uses
client.responses.create()to send a request to the OpenAI API.model: Specifies the model to use (e.g., "gpt-4-turbo-preview").input: The prompt or question for the AI. Note the change frommessagesin the Chat Completions API toinputin the Responses API.
- It prints the AI's response using
response.output.text.
4. Running the Code:
- Execute the Python script from the terminal:
python app.py
Adding a User Interface with Gradio
The video demonstrates how to create a simple user interface for the chatbot using Gradio.
1. Code Modification:
import gradio as gr
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY")
def ask_ai(question):
response = client.responses.create(
model="gpt-4-turbo-preview",
input=question
)
return response.output.text
iface = gr.Interface(
fn=ask_ai,
inputs="text",
outputs="text",
title="My Chatbot"
)
iface.launch()
- Import the
gradiolibrary asgr. - Define a function
ask_ai(question)that takes a question as input, sends it to the OpenAI API, and returns the AI's response. - Create a Gradio interface using
gr.Interface().fn: The function to call when the user submits input (ask_ai).inputs: The type of input (e.g., "text").outputs: The type of output (e.g., "text").title: The title of the chatbot interface.
- Launch the interface using
iface.launch().
2. Running the Code:
- Execute the Python script from the terminal:
python app.py - Gradio will provide a URL to access the chatbot interface in your web browser.
Image Analysis with the Responses API
The video shows how to use the Responses API to analyze images.
1. Code Implementation (image.py):
from openai import OpenAI
import base64
client = OpenAI(api_key="YOUR_API_KEY")
## Function to encode the image
def encode_image(image_path):
with open(image_path, "rb") as image_file:
return base64.b64encode(image_file.read()).decode('utf-8')
## Path to your image
image_path = "path/to/your/image.jpg"
## Getting the base64 string
base64_image = encode_image(image_path)
response = client.responses.create(
model="gpt-4-vision-preview",
input=[
{
"type": "text",
"content": "What teams are playing in this image?"
},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
],
max_tokens=300
)
print(response.output.text)
- The code uses the
gpt-4-vision-previewmodel. - The
inputis a list containing two dictionaries:- One with
type: "text"andcontentcontaining the question about the image. - One with
type: "image_url"andimage_urlcontaining the base64 encoded image.
- One with
- The code then prints the AI's analysis of the image.
2. Running the Code:
- Execute the Python script from the terminal:
python image.py
Utilizing Inbuilt Tools: Web Search and File Search (RAG)
The video explains how to use the Responses API with inbuilt tools like Web Search and File Search.
1. Web Search Tool:
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY")
response = client.responses.create(
model="gpt-4-turbo-preview",
input="Give me two AI news stories from today in two sentences.",
tools=[{"type": "web_search"}]
)
print(response.output.text)
- The
toolsparameter is a list containing a dictionary withtype: "web_search". - The AI will automatically search the web to answer the question.
2. File Search Tool (RAG):
- Upload Data: Upload your custom data (e.g., PDFs) to platform.openai.com under "Storage" and create a vector store.
- Code Implementation (rag.py):
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY")
vector_store_id = "YOUR_VECTOR_STORE_ID"
response = client.responses.create(
model="gpt-4-turbo-preview",
input="Tell me about graph RAG.",
tools=[{"type": "file_search", "database_id": vector_store_id}]
)
print(response.output.text)
- Replace
"YOUR_VECTOR_STORE_ID"with the ID of your vector store. - The
toolsparameter is a list containing a dictionary withtype: "file_search"anddatabase_idset to your vector store ID. - The AI will search the uploaded files to answer the question.
The speaker highlights that this is a simplified implementation of RAG (Retrieval Augmented Generation). The example demonstrates "agentic RAG," where the AI automatically decomposes the query into multiple sub-queries for more effective search.
Creating Custom Tools
The video details how to create and use custom tools with the Responses API. This is crucial for building AI agents.
1. Define the Custom Tool (get_weather function):
import geopy
import openmeteo
def get_weather(location):
geolocator = geopy.Nominatim(user_agent="geoapiExercises")
location_data = geolocator.geocode(location)
latitude = location_data.latitude
longitude = location_data.longitude
openmeteoClient = openmeteo.Client()
weather_data = openmeteoClient.forecast(latitude, longitude, hourly="temperature_2m")
temperature = weather_data.Hourly().Variables(0).ValuesAsNumpy().tolist()[0]
return temperature
- This example defines a
get_weather(location)function that usesgeopyto get the coordinates of a location andopenmeteoto retrieve the current temperature.
2. Define the Tool Definition:
tool_definition = {
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current temperature of a given location.",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The location to get the weather for."
}
},
"required": ["location"]
}
}
}
- This defines the structure and purpose of the
get_weathertool for the AI. - It specifies the tool's name, description, and the parameters it requires (in this case, the
location).
3. Call the Responses API (First Call):
from openai import OpenAI
import json
client = OpenAI(api_key="YOUR_API_KEY")
response = client.responses.create(
model="gpt-4-turbo-preview",
input="What is the weather like in Paris today?",
tools=[tool_definition]
)
print(response.output.text)
- The
toolsparameter now includes thetool_definition. - The AI will recognize that it needs to use the
get_weathertool to answer the question.
4. Extract Tool Call and Parameters:
function_name = response.output.function_call.name
arguments = json.loads(response.output.function_call.arguments)
location = arguments["location"]
- Extract the name of the function to call (
function_name) and its arguments (arguments) from the API response.
5. Execute the Tool:
function_response = get_weather(location)
- Call the
get_weatherfunction with the extracted location.
6. Call the Responses API (Second Call):
response2 = client.responses.create(
model="gpt-4-turbo-preview",
input="What is the weather like in Paris today?",
tools=[tool_definition],
tool_results=[{
"tool_call_id": response.output.function_call.id,
"output": str(function_response)
}]
)
print(response2.output.text)
- Call the API again, providing the original
input, thetool_definition, and thetool_results. - The
tool_resultscontain the output from theget_weatherfunction. - The AI uses this information to generate a natural language response.
The speaker emphasizes that the second API call is necessary to convert the programmatic output of the tool (e.g., a JSON response) into a user-friendly natural language response.
Streaming Responses
The video demonstrates how to stream responses from the Responses API for a more engaging user experience.
1. Code Implementation (stream.py):
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY")
response = client.responses.create(
model="gpt-4-turbo-preview",
input="Write a story with thousand words",
stream=True
)
for chunk in response:
print(chunk.output.text, end="", flush=True)
- Set
stream=Truein theclient.responses.create()call. - Iterate through the
responseobject, which yields chunks of the response as they are generated. - Print each chunk to the console.
LM Studio Integration (Work in Progress)
The video mentions that integration with LM Studio is a work in progress. LM Studio allows users to download and run AI models locally on their computers. The goal is to enable the Responses API to work with local models without requiring an API key or external source. This would provide greater privacy and control over the AI processing.
Nvidia GTC Conference
The video briefly mentions the Nvidia GTC conference, highlighting sessions related to building agentic AI systems and AI agents in production.
Conclusion
The video provides a comprehensive overview of the OpenAI Responses API, highlighting its key features, benefits, and practical applications. It offers step-by-step instructions for creating chatbots, analyzing images, utilizing inbuilt tools, creating custom tools, and streaming responses. The speaker emphasizes the importance of the Responses API for building advanced AI applications and encourages viewers to explore the new features and capabilities. The mention of LM Studio integration suggests a future direction towards more localized and private AI development.
AI summaries can miss context or contain errors. Check important details against the original video.