OpenAI Changed AI Agents Forever (The Update Everyone Missed)

Arseny ShatokhinAbout 5 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • AI Systems: Defined as interconnected elements (reasoning, tool calls, modalities) producing behavior over time, where the whole is greater than the sum of its parts.
  • Elements of a System: Elements, interconnections, and purpose.
  • Feedback Loop: Crucial for system behavior, allowing reflection and correction.
  • Tool Calling: The ability of AI models to use external tools to accomplish tasks.
  • Reasoning: The ability of AI models to think and make decisions based on available information.
  • Modalities: Different forms of input, such as text and images.
  • AGI (Artificial General Intelligence): A system capable of performing most tasks better than humans.
  • Multi-Agent System: A system composed of multiple specialized AI agents working together.

1. Understanding AI Systems

  • Definition of a System: Drawing from Donella Meadows' "Thinking In Systems," a system is a set of interconnected things producing their own behavior over time. The whole is bigger than the sum of its parts. Examples include homes, phones, businesses, and human beings.
  • Components of a System: Elements, interconnections, and purpose. A thermostat is used as an example, with elements like a temperature sensor, AC, and windows, interconnected to maintain a constant temperature through a feedback loop.
  • Feedback Loops: Essential for system behavior. The thermostat example illustrates how the AC kicks in when the temperature rises to maintain a comfortable level.

2. Breakthrough with o4-mini and o3

  • Limitations of Previous Models: Previous large language models lacked a feedback loop, simply taking inputs and producing outputs without reflecting on their actions. This made it difficult to build reliable AI agents.
  • Reasoning Over Tool Calls: Older models were trained to reason over text inputs but not over tool calls. This required adding supervisor agents to reflect on the actions of other agents.
  • o4-mini and o3's Advancement: These models are specifically trained to use tools through reinforcement learning, enabling them to reason about when to use them. This allows them to execute up to 600 tool calls in a sequence.
  • OpenAI Quote: "Were specifically trained to use tools through reinforcement learning, teaching them not just how to use tools, but to reason about when to use them."

3. How o4-mini and o3 Agents Work

  • Standard Agent Workflow: Agents come up with a plan and execute tasks sequentially using tools. If something goes wrong, they cannot correct their course because they can only reason over user inputs, not tool calls.
  • o4-mini/o3 Agent Workflow: These agents can think after every action and select the proper path accordingly. This allows them to automate more complex workflows and correct their course if needed.
  • Reasoning Feedback Loop: The addition of a feedback loop allows the agent to think about each tool call and execute further tools as necessary. This enables the agent to loop by itself without human input, continuing until the task is fully satisfied.
  • Three Elements Utilized: Reasoning, tool calls, and modalities (image and text inputs) are used in a continuous reinforcement feedback loop.

4. Performance and Implications

  • AIME 2025 Benchmark: Demonstrates the impact of adding tools to models. O4-mini with tools achieved 99.5% accuracy, closing the gap significantly compared to O4-mini without tools.
  • Exponential Growth: The release has solved math, which will accelerate the development of new AI systems.
  • Reliability of Agents: Agents can better understand and correct their own mistakes due to reasoning after each tool call.
  • Agent-Driven Insights: Agents can generate novel insights and tell users what to do, reducing the need for long prompts. Users only need to provide access to internal systems, data, and goals.
  • Workflow Automation: Agents can design complex workflows, potentially replacing platforms like Zapier or Make.
  • Agent Evolution: Agents can adjust their own instructions and tools to make workflows more dynamic.
  • Long-Running Agents: Agents can run for days, executing sophisticated tasks without constant user input.

5. Limitations

  • Stupid Mistakes: Models, especially o4-mini, can still make mistakes that humans would not.
  • Hallucinations: O4-mini has a higher rate of hallucinations compared to O3.
  • API Availability: Tool calling is not yet available in the API at the time of recording.

6. Preparing for the Next Generation of AI Agents

  • Take Initiative: Explore and experiment with AI systems, as implementation will soon be outsourced to AI.
  • Start Building Agents Now: Build agents for personal tasks, daily work, or business use cases.
  • Community Invitation: Join the speaker's free community for support, tools, courses, and Q&A calls.
  • Model Selection: Understand the capabilities and costs of different models. O4-mini is recommended for AI agents due to its cost-effectiveness. Use GPT-4.1 for straightforward tasks and O3 for mission-critical tasks with low error tolerance.

7. Is This AGI?

  • Not AGI by Default: The AI systems are not inherently AGI, but they can be made into AGI.
  • Definition of AGI: A system capable of performing most tasks better than humans.
  • Requirements for AGI: Providing the system with necessary tools, knowledge, and instructions.
  • Multi-Agent Systems: Building specialized agents with their own tools and instructions, similar to human organizations.
  • Call to Action: Take initiative and start building agents in specific industries to collectively create AGI.

8. Conclusion

The release of o4-mini and o3 represents a significant advancement in AI systems, particularly in their ability to reason and use tools effectively. While limitations exist, the potential for building reliable, insightful, and long-running AI agents is immense. The key to unlocking this potential lies in taking initiative, building agents now, and understanding the nuances of different models. The speaker believes that by building multi-agent systems, we can collectively create AGI.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.