OpenAI O3: A Breakthrough in Agentic Coding?

Prompt EngineeringAbout 6 min readApr 18, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Agentic Reasoning: The model's ability to independently decide when and how to use tools (like web search) to improve its responses.
  • Sequential Tool Calling/Function Calling: The model's capacity to execute multiple tools or functions in a sequence, analyzing the results of each before proceeding.
  • Multimodal Reasoning: The model's ability to understand and reason about information from different modalities, such as text and images.
  • SDK (Software Development Kit): A set of tools and resources for building applications for a specific platform or technology.
  • Painter's Algorithm: A rendering technique that draws objects from back to front, potentially causing overlapping issues.
  • Model Context Protocol & Agent-to-Agent Protocol: (Likely) Emerging concepts related to how models interact and share information, especially in the context of AI agents.

Model Context Protocol and Agent-to-Agent Protocol

  • The presenter asked O3 about "model context protocol" and "agent-to-agent protocol" without enabling web search.
  • O3 recognized these as potentially new or dynamic terms (especially post-2025) and decided to perform a web search to get accurate definitions.
  • It conducted two passes of web search, first looking up the terms and then refining its search based on the initial results.
  • This behavior is described as "agentic" and contrasts with other models that would simply hallucinate an answer.
  • The presenter compares this to the web search implementation in Clone, which also uses a similar pattern of internal chain of thought and iterative web search queries.

Encyclopedia of Legendary Pokémon

  • The prompt was to create a simple encyclopedia of the first 25 legendary Pokémon, including their types, lore snippets, and images, in a single HTML, CSS, and JavaScript file.
  • The model's internal chain of thought was more agentic than previous models, focusing on understanding the user's requirements and developing a plan.
  • The resulting website was described as one of the best-looking websites generated by an LLM for this prompt, indicating fine-tuning for coding and UI design.
  • When asked to add a search bar, the model provided flawless code while preserving the existing functionality.
  • The presenter emphasizes the importance of requesting complete code from the model, as it sometimes provides only code edits.

Coded TV Channel Generator

  • The prompt was to create a coded TV that lets the user change channels with number keys (0-9). Each channel should have an idea and interesting animation inspired by a classic TV genre, created with p5.js within an 800x800 sketch, without using HTML, and masked to a TV screen area.
  • The model successfully generated different TV channels with animations inspired by genres like classic cartoons, global news, action sports, nature, sci-fi, and cooking.
  • The presenter noted that the animations were not as good as those generated by GPT-4.1 but were still impressive.
  • The presenter then provided a screenshot of the generated weather channel and asked the model to identify any issues based on the original requirements.
  • The model analyzed the image, identified potential layering issues, and suggested debugging strategies, including writing and executing Python code within its chain of thought.
  • The model identified that the weather channel was leaking out of the TV screen mask (though this was not visually apparent) and recommended updates.
  • When the presenter pointed out that the sketch was rectangular instead of square, the model explained that the canvas was created correctly but might be displayed in a way that makes it look rectangular.

Rotating Sphere of ASCII Numbers

  • The prompt was to create a JavaScript simulation of a sphere made of ASCII numbers rotating, with the closest numbers being pure white and the furthest fading to gray on a black background, all in a single line of code.
  • The initial code generated had the gradient reversed (furthest numbers were white, closest were gray).
  • The presenter provided a screenshot of the simulation and asked the model to identify the issue.
  • The model thought for 20 seconds, then for another 47 seconds, before identifying potential Z-sorting issues and the backward painter's algorithm order.
  • The model provided updated code, but it still had the same issue.
  • After another iteration, the model provided a corrected code that solved the problem, demonstrating its ability to learn from feedback and reason over multimodal data.

Bouncing Balls in a Heptagon

  • The prompt was to create a JavaScript animation of 20 balls bouncing inside a heptagon, with specific requirements for ball radius, numbering, starting position, color schemes, and interactions between the balls and the sides of the heptagon, all in a single HTML file.
  • The model successfully generated code that met all the requirements, with the balls starting from the center and bouncing realistically within the heptagon.
  • The presenter noted that the model's code was concise and did not include extensive explanations.

Falling Letters Animation (Failure Case)

  • The prompt was to create a JavaScript animation of falling letters with realistic physics, appearing randomly at the top of the screen with different sizes and falling under Earth's gravity with collision detection.
  • The model generated code that resulted in an error related to an unsupported open type signature (404).
  • Despite several iterations and back-and-forth communication, the model was unable to produce a working animation, demonstrating that it is not a perfect model and can fail in unexpected ways.

Text-to-Image App with Gemini Flash 2.0

  • The prompt was to create a text-to-image app where the user provides a text prompt, and the app uses the Gemini Flash 2.0 native image generation API to create and display the image, with the ability to regenerate and download the image. The presenter asked for everything to be implemented in a single Python file.
  • The presenter intentionally did not provide the API documentation to see how the model would handle the task, given that the Gemini SDK had recently changed.
  • The model's internal chain of thought involved multiple passes of web search to gather information about the Gemini Flash 2.0 image generation API.
  • It decided to use the newer version of the Gemini SDK but initially made a mistake by using configurations from the older SDK.
  • After receiving an error message, the model planned, did a second search, and updated its code.
  • Before running the code, the presenter asked the model to generate an image of how the app was supposed to look, and the resulting image was very close to the actual app that was generated.
  • After a couple of iterations, the app was fully working, self-contained in a single file, and functioned as expected.
  • The presenter was impressed by the model's ability to reason through SDK versions, figure out changes in the SDK, and use sequential tool calling within its chain of thought.

Benchmarks and Performance

  • Independent benchmarks show that O3 performs well compared to Gemini 2.5 Pro, especially on benchmarks highlighted by OpenAI.
  • Gemini 2.5 Pro has a better performance-to-cost ratio than O3.
  • Gemini performs better than O3 and O4-Mini on GPQA (PhD-level question answering) at a much lower cost.
  • O3 is the new standard for code-specific tasks like Ader polyglot, but at a higher cost.

Conclusion

  • O3 is not AGI and makes mistakes, but its ability to use tools in agentic workflows is impressive.
  • Its coding capabilities are state-of-the-art in available benchmarks.
  • The presenter is excited about the potential of O3 and its agentic coding capabilities, especially if integrated into a coding IDE.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.