OpenAI’s Responses API: The Easiest Way to Build a RAG System?

Prompt EngineeringAbout 4 min readMar 19, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Responses API, RAG (Retrieval-Augmented Generation), Vector Store, File Search Tool, Evaluation/Testing Strategy, Recall, Precision, Ragas, LLM-assisted evaluation.

1. Introduction to Responses API

  • OpenAI released new developer tools, including the Responses API.
  • This API is an alternative to the Assistants API (deprecated mid-2026).
  • The Responses API allows the use of tools, differentiating it from the chat completion API.

2. File Search Tool and RAG Implementation

  • The Responses API includes a built-in File Search tool.
  • Users can upload files to OpenAI servers, which creates a vector store.
  • The API retrieves relevant document chunks based on user queries.
  • OpenAI handles chunking, embedding, and retrieval pipelines.
  • Companies are building custom RAG pipelines using this API.
  • Cost: $0.10 per GB of vector storage per day (first GB free), $2.50 per 1000 tool calls.

3. Building a RAG Pipeline with Responses API

  • The video demonstrates building a RAG pipeline using the Responses API.
  • The example uses OpenAI blog posts as the data source.
  • The process includes creating a vector store, uploading files, and querying the store.

3.1. Installation and Setup

  • Install the latest version of the OpenAI Python SDK.
  • Import necessary packages.
  • Store the OpenAI API key securely (e.g., using Google Colab secrets).

3.2. Data Preparation

  • Download relevant documents (e.g., OpenAI blog posts) and store them in a folder.
  • Supported file types include PDFs and others (a detailed list is available).

3.3. Creating and Populating the Vector Store

  • Use client.vector_stores.create to create a vector store, providing a name.
  • The function returns the vector store ID, name, creation timestamp, and file count.
  • A helper function uploads individual files to the vector store using files.create endpoint.
  • Another helper function concurrently uploads all files from a folder to expedite the process.
  • OpenAI handles chunking and embedding after file upload.

3.4. Retrieving Documents from the Vector Store

  • Use the search endpoint on the vector store to retrieve relevant documents based on a query.
  • Provide the vector store ID and the user query.
  • The result is a list of documents with relevance scores.
  • Example: Querying "what is deep research" returns documents from the "introducing deep research" blog post.
  • The system processes documents in terms of characters (approximately 3000 characters per chunk, roughly 800-1000 tokens).

3.5. Generating Responses with the LLM

  • Use the response endpoint for end-to-end retrieval and generation.
  • Provide the user query, the model to use (e.g., gpt-4-turbo-preview), and the tool to use (e.g., file_search).
  • Specify the vector store within the file_search tool.
  • The API retrieves relevant documents and generates a response using the LLM.
  • Example: Querying about "Elon's case" and "deep research" generates responses for both topics.
  • Multiple tools can be combined, such as the web search tool for fallback information retrieval.

4. Evaluation and Testing Strategy

  • Evaluating the RAG pipeline is crucial to measure the quality of responses.
  • Evaluation should occur at both the retrieval and generation steps.
  • Retrieval Step: Measure whether the retrieved documents are relevant to the question.
  • Generation Step: Measure whether the LLM generates appropriate responses based on the retrieved documents.
  • Libraries like Ragas can be used for advanced evaluation (measuring recall, faithfulness, and correctness).

4.1. Creating an Evaluation Dataset

  • Use an LLM to generate questions based on the documents.
  • Prompt: "Can you generate a question that can only be answered using this document?"
  • Example: "What was the outcome of Elon Musk's request for a preliminary injunction against OpenAI?"
  • Human review is essential to validate the generated questions.
  • Paper reference: "Who validates the validators: Aligning LLM-assisted evaluation of LLM outputs with human preferences."

4.2. Evaluating Retrieval Accuracy

  • Loop through the questions and query the vector store.
  • Compare the retrieved document with the original document where the question originated.
  • If the documents match, consider it a pass; otherwise, a fail.
  • Calculate recall and precision at K (e.g., top 5 or top 10 documents).
  • The example uses five questions, resulting in 100% recall.

4.3. Advanced Evaluation Metrics

  • Consider metrics like response relevancy and faithfulness for evaluating generation accuracy.
  • Ragas is recommended for more advanced evaluation metrics.

5. Conclusion

  • The Responses API simplifies the creation of RAG pipelines.
  • It's important to compare the pricing with custom implementations or other solutions.
  • A robust evaluation strategy is essential for measuring the quality of the RAG pipeline.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.