Key Concepts
- OpenAI o1-03: New flagship reasoning model, excels in coding, math, science, visual analysis. High performance but expensive.
- OpenAI o1-04-mini: Compact, cost-efficient model, strong performance for its size, particularly in math, coding, visual reasoning. Outperforms o1-03-mini.
- Autonomous Tool Use: Both models can use tools like web browsing, Python execution, file analysis, image understanding, and image generation agentically.
- Enhanced Reasoning: Models are built to "think longer and reason more deeply" before responding, producing contextual outputs.
- Benchmark Performance: Significant gains on benchmarks like Swaybench, MMLU, AIM, HellaSwag, outperforming previous models and competitors like Gemini 2.5 Pro in several areas.
- Pricing: Specific costs per 1 million tokens provided for input, cached input, and output for both models, highlighting o1-03's higher cost.
- Context Window: Both models feature a 200K context window.
- Practical Testing: o1-03 demonstrated strong capabilities in front-end development, algorithmic coding (Python), SVG generation, math word problems, creative coding (p5.js), scientific paper comprehension, and logical deduction puzzles.
- Future Models: Anticipation for o1-03-pro and GPT-5 (expected July).
- Codex Copilot: New Code Interpreter-like feature mentioned alongside the model launch.
Introduction to New OpenAI Models
OpenAI has launched two new models, o1-03 and o1-04-mini, announced via a tweet by Sam Altman. A new feature similar to Claude's Coder, named Codex Copilot, was also introduced. The o1-03 model is positioned as OpenAI's "smartest and most capable model yet." A key advancement is their ability to autonomously use the full suite of ChatGPT tools, including web browsing, Python execution, file analysis, image understanding, and image generation. These models are designed for deeper reasoning and longer thinking processes, enabling them to use tools agentically and generate highly contextual outputs. The speaker notes that benchmark scores are becoming saturated, and the qualitative reasoning capabilities are becoming more significant differentiators.
OpenAI o1-03: Flagship Reasoning Model
- Capabilities: OpenAI's most powerful reasoning model to date. Excels in coding, math, science, and visual analysis. Sets new state-of-the-art performance on benchmarks like CodeForce, Swaybench, and MMLU. It makes 20% fewer major errors compared to previous models. Particularly strong in programming, business applications, and creative iteration.
- Pricing: The main drawback is its cost.
- Input Tokens: $10 per 1 million tokens
- Cached Input Tokens: $2.50 per 1 million tokens
- Output Tokens: $40 per 1 million tokens
OpenAI o1-04-mini: Compact and Cost-Efficient Model
- Capabilities: A smaller, cost-effective model that delivers performance significantly above its class ("punches way above its weight"). It dominates many benchmarks and outperforms the previous o1-03-mini. Ideal for high-throughput tasks involving math, coding, and visual reasoning.
- Pricing: Offers a much lower cost structure.
- Input Tokens: $1.10 per 1 million tokens
- Cached Input Tokens: $0.275 (27.5 cents) per 1 million tokens
- Output Tokens: $4.40 per 1 million tokens
Anticipated Future Models
The transcript mentions that OpenAI o1-03-pro is expected "sooner than later" and is likely to be a "major leap." Furthermore, GPT-5 is anticipated to be released in July. These future developments suggest continued improvements in pricing, tool use becoming standard, and a rise in agent-based AI systems.
Benchmark Performance Analysis
The new models demonstrate significant improvements across various benchmarks:
- Swaybench (rarified test):
- o1-03: 69.1%
- o1-04-mini: 68.1%
- Gemini 2.5 Pro: 63.8%
- Note: Slightly behind Claude 3.7 in thinking capabilities.
- Math (AIM 2024 & 2025):
- o1-04-mini leads with 93.4% (AIM 2024) and 92.7% (AIM 2025), surpassing both o1-03 and Gemini.
- Reasoning:
- MMLU: o1-03 leads with 82.9%.
- HellaSwag: o1-03 leads with 20.3%.
- GBQA: Gemini 2.5 Pro slightly edges out with 84%.
Overall Assessment: The o1-03 is positioned as a top-tier reasoning and coding model, while the o1-04-mini provides exceptional performance relative to its size and cost. Both represent a "clear leap" over previous generations and current competitors. The speaker speculates these releases might be a response to the unsuccessful or delisted GPT-4.5 launch and are precursors to GPT-5.
Speaker's Analysis and Recommendation
Despite both models having a 200K context window, the speaker argues against using the expensive o1-03 for coding tasks.
- Argument: The o1-04-mini is "nearly as good" for coding, "way cheaper," and delivers "excellent performance."
- Reasoning: As a reasoning model, o1-03 takes multiple steps, which can quickly consume budget.
- Recommendation: "I think it might be smarter to use the O4 mini for coding cuz it is going to deliver the same sort of performance and it's going to be way cheaper than the 03." Use o1-03 primarily for complex reasoning tasks.
Practical Demonstrations: Testing o1-03
The speaker tested the o1-03 model (accessed via ChatGPT Pro plan; API access also available) on several prompts:
- Modern Note-Taking App Front-End:
- Goal: Assess UI design and web development skills.
- Outcome: First iteration was functional. The second iteration was significantly improved with better aesthetics, animations, export, clear, search, and dark mode features.
- Result: Pass.
- Game of Life (Python):
- Goal: Assess algorithmic design (2D/3D arrays, loops, input handling).
- Outcome: Generated a functional Python script for the Game of Life runnable in the terminal.
- Result: Pass.
- Symmetrical SVG Butterfly:
- Goal: Test spatial reasoning, symmetry logic, SVG syntax knowledge.
- Outcome: Produced a "beautiful" SVG representation, considered one of the best generations the speaker has seen for this prompt, although the wings weren't explicitly designed as hoped.
- Result: Pass.
- Train Meeting Time Problem:
- Goal: Assess math word problem understanding (distance, speed, time).
- Outcome: Correctly calculated the meeting time (1:12 p.m.) and showed clear steps. The model's speed was noted positively.
- Result: Pass.
- TV Simulation (p5.js):
- Goal: Assess creative coding, animation, and understanding of interactive programming (p5.js).
- Outcome: Generated a "beautiful TV app" with nine distinct channels showing different simulations. Considered the "best generation" received for this prompt, despite one page not rendering correctly.
- Result: Pass.
- Climate Modeling Paper Analysis:
- Goal: Assess reasoning, comprehension, and summarizing scientific findings from provided text sections.
- Outcome: Accurately explained why the hybrid deep ensemble model outperformed traditional models, focusing on the correct points and advantages in arid regions based on the provided text.
- Result: Pass.
- Detective Puzzle:
- Goal: Assess logical deduction, identifying contradictions, and truth-finding based on constraints.
- Outcome: Correctly identified the guilty party (David) by logically analyzing conflicting statements from five suspects, assuming each was guilty in turn to find the contradiction.
- Result: Pass.
Overall Test Impression: The o1-03 model proved "exceptional" across coding, reasoning, and creative tasks. However, the speaker reiterates the cost concern for o1-03, especially for coding, suggesting o1-04-mini as a more practical choice for such tasks, while reserving o1-03 for demanding reasoning challenges. It's becoming increasingly difficult to create prompts that truly challenge these advanced models.
Access and Availability
The o1-03 and o1-04-mini models are currently accessible through ChatGPT Pro plans. They are expected to become available later for free tiers, likely with rate limits. API access is also available for developers.
Conclusion/Synthesis
OpenAI's release of o1-03 and o1-04-mini marks a significant step forward in AI capabilities. o1-03 offers state-of-the-art reasoning and broad task proficiency but comes at a high price point. o1-04-mini provides remarkable performance, especially in coding and math, at a much lower cost, making it a highly practical option for many applications. Both models feature advanced autonomous tool use and strong benchmark results, surpassing previous models and competitors in key areas. Practical tests confirm o1-03's impressive abilities across diverse domains like coding, logic, and comprehension. The speaker recommends leveraging o1-04-mini for cost-sensitive tasks like coding, while utilizing o1-03 for its superior reasoning power where needed. These releases signal OpenAI's continued rapid advancement and build anticipation for upcoming models like o1-03-pro and GPT-5.
AI summaries can miss context or contain errors. Check important details against the original video.





