Key Concepts
- Responses API: OpenAI's new API intended to supersede the Chat Completions API.
- Chat Completions API: The existing API for interacting with OpenAI models, primarily focused on chat-based interactions.
- Backward Compatibility: The Responses API is a superset of the Chat Completions API, meaning it supports everything the latter does, plus more.
- Migration Timeline: OpenAI plans to sunset the Chat Completions API by the end of 2026.
- Developer Role: A new role in the Responses API, alongside system, user, and assistant, allowing for more granular control over model behavior.
- Conversation State Management: A new feature that allows the API to track conversation history using response IDs, simplifying context management.
- Function Calling: Remains largely the same as in the Chat Completions API.
- Structured Output: Obtaining data from the model in a structured format (e.g., JSON), supported through JSON schema and Pydantic models.
- Web Search: A built-in tool that allows the model to browse the web for information.
- File Search: A built-in tool that allows the model to search through uploaded files.
- Reasoning Models: Improved support for reasoning models with a new "effort" parameter.
- Agent SDK: OpenAI's new Agent SDK, intended to replace Swarm, for building AI agents.
Responses API: Introduction and Key Differences
The Responses API is OpenAI's new API, designed to be a superset of the Chat Completions API. This means that everything achievable with the Chat Completions API is also possible with the Responses API, along with additional features. OpenAI intends to sunset the Chat Completions API by the end of 2026, making the Responses API the new standard.
- Backward Compatibility: The Responses API is a superset of the Chat Completions API.
- Migration Timeline: Chat Completions API will be sunsetted by the end of 2026.
- Simplified Interface: The Responses API simplifies the interface for different types of interactions, moving beyond just chat.
- New Features: Native support for web search, a new developer role, improved support for reasoning models, built-in file and factor search, and simplified conversation state management.
Example:
- Chat Completions API:
client.chat.completions.create(...) - Responses API:
client.responses.create(input="Write a one sentence bedtime story", model="gpt-4o")
The Responses API simplifies the process by using an input parameter instead of the more verbose messages object. The output is also simplified, with direct access to the text via response.output.text.
New Features in Detail
Text Prompting and the Developer Role
The Responses API introduces a new "developer" role, allowing for more granular control over model behavior. This role can be used in two ways:
- Instructions Parameter: Setting the
instructionsparameter in the API call.client.responses.create(input="Question about semicolons in JavaScript", model="gpt-4o", instructions="Talk Like a Pirate")
- Developer Role in Messages: Similar to the existing
system,user, andassistantroles.client.responses.create(input=[{"role": "developer", "content": "Talk Like a Pirate"}, {"role": "user", "content": "Question about semicolons in JavaScript"}], model="gpt-4o")
Chain of Command: OpenAI has defined a hierarchy for how different roles and instructions are processed: Platform > System > Developer > User. This means that a system prompt can override a developer message.
Example:
- System Prompt: "Talk Like a Pirate"
- Developer Prompt: "Don't Talk Like a Pirate"
In this case, the system prompt will override the developer prompt, and the model will talk like a pirate.
Conversation State Management
The Responses API simplifies conversation state management by allowing you to reference previous responses using their IDs.
- Previous Approach (Chat Completions API): Manually managing message history and passing the entire context with each API call.
- New Approach (Responses API): Using the
previous_response_idparameter to reference a previous interaction.
Example:
- First API call:
client.responses.create(input="Tell me a joke", model="gpt-4o") - Second API call:
client.responses.create(input="Explain why this is funny", model="gpt-4o", previous_response_id=response.id)
By default, the store parameter is set to True, meaning that the conversation history is stored on the OpenAI platform. Setting store to False will prevent this.
Function Calling
Function calling remains largely the same as in the Chat Completions API. You specify functions in the form of tools and pass them to the API.
Example:
tools = [{"type": "function", "function": {"name": "send_email", ...}}]
client.responses.create(input="Can you send an email to Elon and Kya?", model="gpt-4o", tools=tools)
Structured Output
The Responses API supports structured output in two ways:
- JSON Schema: Specifying a JSON schema to define the desired output format.
- Pydantic Models: Using Pydantic models to define the output structure.
Example (JSON Schema):
text_format = {"type": "calendar_event", "properties": {"name": ..., "date": ..., "participants": ...}}
client.responses.create(input="Extract event details from text", model="gpt-4o", text_format=text_format)
Example (Pydantic Models):
class CalendarEvent(BaseModel):
name: str
date: str
participants: List[str]
client.responses.create(input="Extract event details from text", model="gpt-4o", model=CalendarEvent)
Using Pydantic models allows you to directly access the structured data as a Pydantic object.
Web Search
The Responses API includes a built-in web search tool.
Example:
tools = [{"type": "web_search_preview"}]
client.responses.create(input="What are the best restaurants around the Dam in Amsterdam?", model="gpt-4o", tools=tools)
You can also specify the user location to get more relevant search results.
File Search
The Responses API allows you to upload files and perform semantic search on them.
Process:
- Upload File: Upload a file to the OpenAI platform.
- Create Vector Store (Knowledge Base): Create a vector store to store the file's embeddings.
- Add File to Vector Store: Add the uploaded file to the vector store.
- Perform Search: Use the Responses API to perform a semantic search on the vector store.
Example:
tools = [{"type": "file_search", "vector_store_id": "your_vector_store_id"}]
client.responses.create(input="What is deep research by OpenAI?", model="gpt-4o", tools=tools)
You can control the number of results returned by setting the max_num_results parameter.
Reasoning Models
The Responses API provides improved support for reasoning models with a new "effort" parameter.
Example:
client.responses.create(input="Solve this complex problem", model="gpt-4o", reasoning={"effort": "high"})
The effort parameter can be set to "low", "medium", or "high", controlling the amount of reasoning performed by the model.
Agent SDK
OpenAI has also announced a new Agent SDK, intended to replace Swarm, for building AI agents. This SDK is beyond the scope of this video.
Conclusion
The Responses API is a significant update from OpenAI, offering a simplified interface and new features for building AI applications. While these features can make development easier, it's important to be aware of the potential drawbacks of offloading too much logic to the OpenAI API. Maintaining control over context and information flow is crucial for building robust and scalable AI applications.
AI summaries can miss context or contain errors. Check important details against the original video.





