GPT-5 (General, Mini & Nano) Fully Tested: The Worst Models of 2025. I am so disappointed.

AICodeKingAbout 4 min readAug 8, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • GPT5, GPT5 Mini, GPT5 Nano: New GPT models released via API.
  • High Reasoning: The best version of the tested models.
  • Benchmark Testing: A standardized set of prompts used to evaluate model performance.
  • Synthetic Data: Data generated by a model to distill its capabilities into a smaller model.
  • Code Rendering: The ability of a model to generate code that produces a visual output.
  • Markdown Formatting: A way to format text, including code blocks, for readability.
  • AGI (Artificial General Intelligence): A hypothetical level of AI that can perform any intellectual task that a human being can.

GPT5 Model Testing and Results

  • Overall Performance: The GPT5 model performed poorly on the benchmark tests, significantly underperforming compared to previous models and competitors.
  • Specific Failures:
    • Out of 10 questions, GPT5 only answered one correctly (a riddle, the easiest question).
    • Failed at math problems.
    • Failed to render a 3D floor plan, even after multiple attempts and confirmation tests on Open Router and T3 chat. The generated code consistently failed to produce a visual output.
    • The model's output formatting was flawed, requiring a system prompt modification to enforce markdown formatting with code blocks.
    • Generated a poor SVG image of a panda, with the burger appearing inside the panda's stomach instead of in its hand.
    • Failed to render a Pokeball in 3JS.
    • Failed to render a chessboard with autoplay functionality, despite a previous model (Horizon Beta) passing this test.
    • Failed to create a web version of Minecraft.
    • Failed to create a butterfly flying in the garden.
    • Failed to create a CLI tool for image conversion in Rust.
  • Consistency Across Models: The GPT5 Mini and Nano models performed similarly poorly, with only the riddle question being answered correctly.
  • Benchmark Ranking: The GPT5 models scored in the 10% category, on par with GPoss and similar models, far below Sonnet (with max thinking) and Opus 4.1.

Comparison with Other Models

  • Sonnet and Opus 4.1: These models remain superior to GPT5 in the benchmark tests. Opus has better code quality than Sonnet, but Sonnet scores higher overall.
  • Gemini 2.5 Pro: This model is considered good at front-end tasks but fails at logical reasoning tasks requiring precise dimensions and specifications (e.g., 3JS).
  • Gemini Flash: The presenter would rather use Gemini Flash than GPT5.

Possible Explanations for Poor Performance

  • Synthetic Data Training: The presenter suggests that GPT5 may have been trained primarily on synthetic data, which can lead to good performance on benchmark questions (which are easily compressed) but poor performance on real-world tasks.
  • Microsoft FI Models: The presenter draws a parallel to Microsoft FI models, which also rely heavily on synthetic data.
  • Regression in Performance: The presenter expresses surprise that GPT5 is significantly worse than previous models like GPT4 Mini.

Presenter's Opinion and Recommendations

  • Disappointment: The presenter is highly disappointed with the GPT5 models and considers them "amazingly worse" than existing alternatives.
  • Lack of Real-World Usefulness: The presenter believes that the GPT5 models are not suitable for real-world use cases.
  • Preference for Anthropic and Gemini: The presenter's respect for models from Anthropic and Gemini has increased due to the poor performance of GPT5. The presenter is willing to pay for Anthropic and Gemini models.
  • Call to Action: The presenter encourages viewers to test the GPT5 models themselves and share their experiences in the comments.

Dart (Sponsor)

  • AI-Powered Project Management: Dart combines traditional project management with AI features.
  • AI Capabilities: Dart's AI can brainstorm project ideas, generate task lists, and complete assignments.
  • Custom Agents: Users can create custom agents that trigger from built-in integrations, N8N workflows, or custom web hooks. Examples include coding agents, marketing agents, and mailing agents.
  • Integration: Dart integrates with existing workflows through its MCP server, connecting to Claude, Chat, GPT, and other AI tools.
  • Pricing: Most features are free, with premium options starting at $8 per month.

Conclusion

The new GPT5 models (GPT5, GPT5 Mini, and GPT5 Nano) performed surprisingly poorly in benchmark tests, significantly underperforming compared to previous models and competitors like Sonnet, Opus, and Gemini. The presenter attributes this poor performance to potential over-reliance on synthetic data training, leading to good performance on compressed benchmark questions but failure in real-world applications. The presenter expresses strong disappointment and recommends using alternative models from Anthropic and Gemini instead.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.