OpenAI Responses API: EASY Beginners Tutorial in 14 Mins!

Mervin PraisonAbout 7 min readMar 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Responses API, Chat Completions API, File Search, RAG (Retrieval Augmented Generation), Computer Use, Code Interpreter, Custom Tools, Web Search Tool, Streaming, OpenAI Platform, Gradio, LM Studio, Agentic AI Systems.

Responses API: A Powerful Tool Replacing Chat Completions

OpenAI's Responses API is presented as a significant upgrade to the Chat Completions API, designed to power a wider range of AI applications. The key advantage of the Responses API is its support for features not available in the older Chat Completions API, including:

  • File Search: Enables the AI to search through uploaded files (e.g., PDFs) to answer questions, facilitating Retrieval Augmented Generation (RAG).
  • Computer Use: Allows the AI to control the user's computer, automating repetitive tasks.
  • Code Interpreter: (Coming Soon) Will enable the AI to generate and execute code.

The speaker emphasizes that understanding the Responses API is crucial for anyone building AI applications.

Basic Chatbot Creation

The video provides a step-by-step guide to creating a basic chatbot using the Responses API.

1. Installation:

  • Install necessary Python packages using pip:
    • pip install openai geopy gradio
    • openai: The core OpenAI library.
    • geopy: Used for geolocation in custom tools.
    • gradio: Used for creating a user interface.

2. API Key Setup:

  • Obtain an OpenAI API key from platform.openai.com.
  • Set the API key in your code.

3. Code Implementation (app.py):

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY")

response = client.responses.create(
    model="gpt-4-turbo-preview",
    input="Write a one sentence bedtime story about a unicorn."
)

print(response.output.text)
  • The code imports the OpenAI library.
  • It initializes an OpenAI client with your API key.
  • It uses client.responses.create() to send a request to the OpenAI API.
    • model: Specifies the model to use (e.g., "gpt-4-turbo-preview").
    • input: The prompt or question for the AI. Note the change from messages in the Chat Completions API to input in the Responses API.
  • It prints the AI's response using response.output.text.

4. Running the Code:

  • Execute the Python script from the terminal: python app.py

Adding a User Interface with Gradio

The video demonstrates how to create a simple user interface for the chatbot using Gradio.

1. Code Modification:

import gradio as gr
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY")

def ask_ai(question):
    response = client.responses.create(
        model="gpt-4-turbo-preview",
        input=question
    )
    return response.output.text

iface = gr.Interface(
    fn=ask_ai,
    inputs="text",
    outputs="text",
    title="My Chatbot"
)

iface.launch()
  • Import the gradio library as gr.
  • Define a function ask_ai(question) that takes a question as input, sends it to the OpenAI API, and returns the AI's response.
  • Create a Gradio interface using gr.Interface().
    • fn: The function to call when the user submits input (ask_ai).
    • inputs: The type of input (e.g., "text").
    • outputs: The type of output (e.g., "text").
    • title: The title of the chatbot interface.
  • Launch the interface using iface.launch().

2. Running the Code:

  • Execute the Python script from the terminal: python app.py
  • Gradio will provide a URL to access the chatbot interface in your web browser.

Image Analysis with the Responses API

The video shows how to use the Responses API to analyze images.

1. Code Implementation (image.py):

from openai import OpenAI
import base64

client = OpenAI(api_key="YOUR_API_KEY")

## Function to encode the image
def encode_image(image_path):
  with open(image_path, "rb") as image_file:
    return base64.b64encode(image_file.read()).decode('utf-8')

## Path to your image
image_path = "path/to/your/image.jpg"

## Getting the base64 string
base64_image = encode_image(image_path)

response = client.responses.create(
    model="gpt-4-vision-preview",
    input=[
        {
            "type": "text",
            "content": "What teams are playing in this image?"
        },
        {
            "type": "image_url",
            "image_url": {
                "url": f"data:image/jpeg;base64,{base64_image}"
            }
        }
    ],
    max_tokens=300
)

print(response.output.text)
  • The code uses the gpt-4-vision-preview model.
  • The input is a list containing two dictionaries:
    • One with type: "text" and content containing the question about the image.
    • One with type: "image_url" and image_url containing the base64 encoded image.
  • The code then prints the AI's analysis of the image.

2. Running the Code:

  • Execute the Python script from the terminal: python image.py

Utilizing Inbuilt Tools: Web Search and File Search (RAG)

The video explains how to use the Responses API with inbuilt tools like Web Search and File Search.

1. Web Search Tool:

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY")

response = client.responses.create(
    model="gpt-4-turbo-preview",
    input="Give me two AI news stories from today in two sentences.",
    tools=[{"type": "web_search"}]
)

print(response.output.text)
  • The tools parameter is a list containing a dictionary with type: "web_search".
  • The AI will automatically search the web to answer the question.

2. File Search Tool (RAG):

  • Upload Data: Upload your custom data (e.g., PDFs) to platform.openai.com under "Storage" and create a vector store.
  • Code Implementation (rag.py):
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY")

vector_store_id = "YOUR_VECTOR_STORE_ID"

response = client.responses.create(
    model="gpt-4-turbo-preview",
    input="Tell me about graph RAG.",
    tools=[{"type": "file_search", "database_id": vector_store_id}]
)

print(response.output.text)
  • Replace "YOUR_VECTOR_STORE_ID" with the ID of your vector store.
  • The tools parameter is a list containing a dictionary with type: "file_search" and database_id set to your vector store ID.
  • The AI will search the uploaded files to answer the question.

The speaker highlights that this is a simplified implementation of RAG (Retrieval Augmented Generation). The example demonstrates "agentic RAG," where the AI automatically decomposes the query into multiple sub-queries for more effective search.

Creating Custom Tools

The video details how to create and use custom tools with the Responses API. This is crucial for building AI agents.

1. Define the Custom Tool (get_weather function):

import geopy
import openmeteo

def get_weather(location):
    geolocator = geopy.Nominatim(user_agent="geoapiExercises")
    location_data = geolocator.geocode(location)
    latitude = location_data.latitude
    longitude = location_data.longitude

    openmeteoClient = openmeteo.Client()
    weather_data = openmeteoClient.forecast(latitude, longitude, hourly="temperature_2m")
    temperature = weather_data.Hourly().Variables(0).ValuesAsNumpy().tolist()[0]
    return temperature
  • This example defines a get_weather(location) function that uses geopy to get the coordinates of a location and openmeteo to retrieve the current temperature.

2. Define the Tool Definition:

tool_definition = {
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current temperature of a given location.",
        "parameters": {
            "type": "object",
            "properties": {
                "location": {
                    "type": "string",
                    "description": "The location to get the weather for."
                }
            },
            "required": ["location"]
        }
    }
}
  • This defines the structure and purpose of the get_weather tool for the AI.
  • It specifies the tool's name, description, and the parameters it requires (in this case, the location).

3. Call the Responses API (First Call):

from openai import OpenAI
import json

client = OpenAI(api_key="YOUR_API_KEY")

response = client.responses.create(
    model="gpt-4-turbo-preview",
    input="What is the weather like in Paris today?",
    tools=[tool_definition]
)

print(response.output.text)
  • The tools parameter now includes the tool_definition.
  • The AI will recognize that it needs to use the get_weather tool to answer the question.

4. Extract Tool Call and Parameters:

function_name = response.output.function_call.name
arguments = json.loads(response.output.function_call.arguments)
location = arguments["location"]
  • Extract the name of the function to call (function_name) and its arguments (arguments) from the API response.

5. Execute the Tool:

function_response = get_weather(location)
  • Call the get_weather function with the extracted location.

6. Call the Responses API (Second Call):

response2 = client.responses.create(
    model="gpt-4-turbo-preview",
    input="What is the weather like in Paris today?",
    tools=[tool_definition],
    tool_results=[{
        "tool_call_id": response.output.function_call.id,
        "output": str(function_response)
    }]
)

print(response2.output.text)
  • Call the API again, providing the original input, the tool_definition, and the tool_results.
  • The tool_results contain the output from the get_weather function.
  • The AI uses this information to generate a natural language response.

The speaker emphasizes that the second API call is necessary to convert the programmatic output of the tool (e.g., a JSON response) into a user-friendly natural language response.

Streaming Responses

The video demonstrates how to stream responses from the Responses API for a more engaging user experience.

1. Code Implementation (stream.py):

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY")

response = client.responses.create(
    model="gpt-4-turbo-preview",
    input="Write a story with thousand words",
    stream=True
)

for chunk in response:
    print(chunk.output.text, end="", flush=True)
  • Set stream=True in the client.responses.create() call.
  • Iterate through the response object, which yields chunks of the response as they are generated.
  • Print each chunk to the console.

LM Studio Integration (Work in Progress)

The video mentions that integration with LM Studio is a work in progress. LM Studio allows users to download and run AI models locally on their computers. The goal is to enable the Responses API to work with local models without requiring an API key or external source. This would provide greater privacy and control over the AI processing.

Nvidia GTC Conference

The video briefly mentions the Nvidia GTC conference, highlighting sessions related to building agentic AI systems and AI agents in production.

Conclusion

The video provides a comprehensive overview of the OpenAI Responses API, highlighting its key features, benefits, and practical applications. It offers step-by-step instructions for creating chatbots, analyzing images, utilizing inbuilt tools, creating custom tools, and streaming responses. The speaker emphasizes the importance of the Responses API for building advanced AI applications and encourages viewers to explore the new features and capabilities. The mention of LM Studio integration suggests a future direction towards more localized and private AI development.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.