OpenAI's New Reasoning Models: 03 and 04 Mini - A Detailed Summary
Key Concepts:
- Reasoning Models: AI models capable of complex problem-solving, coding, math, science, and visual perception.
- Tool Usage: The ability of a model to effectively utilize external tools or functions (e.g., web search, Python interpreter, image generation) to enhance its capabilities.
- Agentic Use: Applying AI models in autonomous systems or agents that can perform tasks independently, leveraging reasoning and tool usage.
- Multimodal Reasoning: The ability of a model to process and reason with multiple types of data, such as text and images.
- Reinforcement Learning (RL): A machine learning technique used to train models to make decisions in an environment to maximize a reward.
- Codeex CLI: OpenAI's open-source command-line interface for interacting with reasoning models.
- Instruction Following: The ability of a model to accurately and consistently adhere to user instructions.
- Benchmarks: Standardized tests used to evaluate the performance of AI models.
- Scaling: Increasing the computational resources (e.g., compute, data) used to train or run AI models.
Model Overview: 03 and 04 Mini
OpenAI has released two new reasoning models, 03 and 04 Mini, designed to replace the original 01 models. These models are specifically designed for agentic use cases and boast significant improvements in tool usage and multimodal reasoning.
- 03: The most powerful reasoning model, excelling in coding, math, science, and visual perception. It demonstrates a significant performance boost on multimodal benchmarks, enabling it to analyze images, charts, and graphs. External evaluations show 20% fewer major errors compared to the 01 model on real-world tasks.
- 04 Mini: A smaller, cost-optimized model designed for fast and efficient reasoning, particularly in code, mathematics, and visual tasks. It offers remarkable performance for its size and cost, outperforming the existing 03 Mini on certain benchmarks at a similar cost.
Key Improvements and Capabilities
-
Tool Usage:
- The models can now effectively use tools, a major limitation of previous reasoning models.
- They can access web search, analyze uploaded files, and generate images (potential feature).
- Reinforcement learning was used to train the models not just how to use tools, but to reason about when to use them.
- The models can reason through the output of tools and modify their plan accordingly.
-
Multimodal Reasoning:
- The models can analyze images, charts, and graphs.
- They can integrate images directly into their chain of thought, enabling problem-solving that blends visual and textual reasoning.
- This unlocks new applications, particularly in industries like manufacturing, where models can visually reason through products.
-
Instruction Following:
- Improved instruction following capabilities are a key focus, ensuring the models adhere to user instructions.
- The models demonstrate state-of-the-art performance on instruction following benchmarks, specifically with tool usage.
-
Coding Performance:
- The models achieve state-of-the-art performance on coding benchmarks, surpassing Gemini 2.5 Pro on some tasks.
- They excel in code editing and generation, demonstrating significant improvements over previous models.
- However, real-world coding performance on larger projects still needs to be tested.
Benchmarks and Performance
- The models have saturated many existing coding and math benchmarks, indicating a need for more challenging evaluations.
- On the AIM 2025 competition, 04 Mini achieves 99.5% accuracy with a Python interpreter.
- Significant performance improvements are observed on benchmarks like Sweep Bench Verified (03: 69%, 04 Mini: 68%) and AD Polyglot Code Editing (03: 81% on high settings).
- The models demonstrate state-of-the-art results on the Top Bench function calling benchmark.
- The blog post compares the new models only to previous OpenAI models, and the speaker suggests that including competitor models on the same plots would provide more transparency.
Reinforcement Learning and Scaling
- Large-scale reinforcement learning played a crucial role in the development of the models.
- Increased compute during training and inference leads to better performance, suggesting that scaling opportunities remain.
- The models were trained to reason about when to use specific tools through reinforcement learning.
Codex CLI
- OpenAI introduced Codex CLI, an open-source command-line interface for interacting with the reasoning models.
- It allows users to run the models on their local machines and reason through images.
- It is positioned as a competitor to cloud code.
Pricing and Availability
- The models are available through the API and within ChatGPT.
- 04 Mini is less expensive than 4.1, even though its benchmarks look better.
- 03 is much lower cost compared to the existing 01 models.
- 04 Mini will have much higher usage limits than 03.
Notable Quotes
- "For the first time these models can integrate images directly into their chain of thought. They don't just see an image, they think with it." - Highlighting the multimodal reasoning capabilities.
Conclusion
OpenAI's release of the 03 and 04 Mini reasoning models represents a significant advancement in AI capabilities, particularly in tool usage, multimodal reasoning, and instruction following. The models demonstrate state-of-the-art performance on various benchmarks and offer cost-effective solutions for agentic applications. The introduction of Codex CLI further enhances accessibility and usability. These advancements set a high bar for competitors like Google and Claude, driving further innovation in the field. The speaker expresses excitement about experimenting with the models and providing more detailed insights in future videos.
AI summaries can miss context or contain errors. Check important details against the original video.





