Yep, o3-mini is WORTH the money - Build your own reasoning agent

AI JasonAbout 4 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Reasoning Models: Language models like O3 Mini and DeepSeek R1 that excel at complex planning, synthesis, and multi-step tasks.
  • Tool Calling Agents: Agents that can use external tools (e.g., web scrapers, search engines) to gather information and perform actions.
  • Deep Research Agent: An agent designed for in-depth research, capable of multiple rounds of actions, reflection, and synthesis.
  • GIA Benchmark: A benchmark dataset from Meta for evaluating general AI assistant tasks, including research and summarization.
  • Function Calling/Structured Output: The ability of a language model to output structured data, specifying which function to call and with what arguments.
  • Scrippers/Data Scrapers: Tools used to extract data from websites.
  • Reflection: A process where the agent analyzes its previous actions and adjusts its strategy.

O3 Mini Model and Deep Research Agents

Introduction to O3 Mini

  • The O3 Mini model is highlighted as an underhyped but significant release from OpenAI.
  • It supports function calling and structured output, making it suitable for building production-ready agents.
  • It's significantly cheaper (93% less) and faster (4x) than the 01 model.
  • OpenAI has released a fine-tuned version of O3 Mini called "deep research mode" in ChatGPT, specifically designed for research agents.

Deep Research Agent Capabilities

  • Deep research agents can perform multiple rounds of research actions, including reflection and synthesis.
  • This allows them to conduct deeper research compared to agents based on standard models like GPT-4.
  • The speaker suggests that reasoning models will significantly impact the AI landscape by 2025.

Performance Advantages of Reasoning Models

  • Reasoning models can overcome hurdles faced by standard language models in complex tasks.
  • The O3 Mini model is even cheaper than GPT-4, making it an attractive option for autonomous agent systems.

Replicating Deep Research Agent

Objective

  • The video aims to demonstrate how to replicate a deep research agent using reasoning models like O3 Mini or DeepSeek R1.
  • It includes a side-by-side comparison of the same research task between reasoning agents and standard agents.

Methodology

  • Two research agents are created: one using a reasoning model (O3 Mini) and the other using a standard model (GPT-4).
  • These agents are tested using tasks from the GIA benchmark.
  • The GIA benchmark includes tasks that require research, synthesis, and summarization abilities.
  • Each task has a clear, unambiguous answer for verification.

Agent Architecture

  • The research agents are built as basic tool-calling agents.
  • Tool calling allows the agent to output actions and inputs for external tools.
  • The agent's system prompt defines its role as an internet researcher who uses tools to find the latest data.
  • The reasoning effort is set to "high" for the O3 Mini agent, balancing speed and accuracy.

Agent Workflow

  1. The language model generates an output.
  2. If the output contains a tool call, the function name and inputs are extracted.
  3. The corresponding action function is called, and the result is inserted into the conversation history.
  4. This process repeats until the agent believes the task is finished and outputs a final response.

Tools and Schemas

  • The agent has access to tools like website scrapers (using platforms like FileCor and SpiderCloud) and Google Search (using SerpApi).
  • A "reflection" tool is included to allow the model to analyze its progress before providing the final answer.
  • Tool schemas are used to communicate to the language model when to use each tool and what inputs are required.
  • OpenAI's playground can be used to generate these schemas automatically.

Experimental Results and Comparison

Research Task 1: Invasive Species

  • Task: Identify the five-digit zip codes where a fish species popularized by the movie "Finding Nemo" was found as a non-native species before 2020, according to USGS.
  • The GPT-4 agent failed to find the zip code information.
  • The O3 Mini-based agent successfully completed the task.

Research Task 2: Nano Compound Study

  • Task: Identify the nano compound studied in a natural journalist scientific reports conference proceeding from 20 in an article that didn't mention plasmon or plasmonic.
  • The GPT-4 agent provided the wrong answer.
  • The O3 Mini agent provided the correct answer.
  • Interestingly, even the deep research function in ChatGPT struggled with this task.

General Findings

  • O3 Mini-based agents generally perform better than GPT-4 agents on tasks requiring significant planning and synthesis.
  • There were cases where both agents failed, but the deep research function in ChatGPT succeeded.
  • ChatGPT seems to have access to better scraping tools, allowing it to access websites that the standard scrapers couldn't.

Conclusion and Future Directions

Key Takeaway

  • If your agent task requires complex planning and synthesis, consider using the O3 Mini model to potentially improve performance.

Future Development

  • Building specialized data scrapers for specific vertical research topics can further enhance the performance of research agents.

Resources

  • The notebook used in the video will be available in the AI Builder Club community.
  • The community also provides tutorials on deploying agents as API services.

Final Statement

  • The speaker believes that this is just the beginning of a new era of agents utilizing reasoning capabilities.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.