GPT-5 vs Claude vs Gemini: 7 brutal real-life tests

Authority Hacker PodcastAbout 5 min readAug 10, 2025Watch original
THE SUMMARYAI-generated

GPT5 Massive Upgrade: A Deep Dive for Online Business Owners

Key Concepts:

  • GPT5 and GPT5 Mini: New models from OpenAI, offering improved performance and affordability.
  • Hallucinations: Instances where the model provides incorrect or fabricated information.
  • API Pricing: The cost of accessing the models programmatically, a key factor for developers.
  • Benchmarks vs. Vibe: Objective performance metrics versus subjective user experience.
  • Context Window: The amount of text the model can consider at once, impacting complex tasks.
  • AISMs: Characteristics that make AI-generated content easily identifiable.

1. Overview of GPT5 Release

  • Massive Upgrade: The release of the GPT5 family of models is a significant upgrade for all users, especially online business owners.
  • Affordable API: OpenAI has made these models very affordable in the API space, beating pretty much anything else.
  • Hallucination Issue: Despite improvements, GPT5 still hallucinates more than expected, with examples of deceptive responses.
  • Game Changer for Normal Users: GPT5 simplifies the user experience by removing the need to select specific models (e.g., GPT-4, GPT-3). The system automatically uses the best model for the task.
  • Reasoning Mode: GPT5 has a "thinking" mode that is triggered for complex questions, providing deeper analysis and better answers.
  • Usage Limits: The normal GPT5 has a limit of 80 messages per hour, while the "thinking" mode has a limit of 200 messages per week.
  • Free Access: GPT5 is available to everyone, even without a paid subscription, but with a smaller context window and no "thinking" mode.

2. Performance and Benchmarks

  • Crushing Benchmarks: GPT5 is crushing everyone in benchmarks, including Gemini 2.5 Pro.
  • Top Performance: GPT5 excels in text, webdev, and vision tasks.
  • Superior Intelligence: Even the medium version of GPT5 surpasses other models in intelligence.
  • Good Vibe: The model is user-friendly, follows instructions well, and does exactly what is asked.

3. Hallucinations and Integrations

  • Hallucination Examples: The presenter encountered two instances of significant hallucinations in five chats.
    • MCP Example: GPT5 provided a detailed but entirely fabricated guide on connecting MCPs to a Chat GPT account, including non-existent file configurations.
  • Limited Integrations: Chat GPT with GPT5 feels like a walled garden, lacking connections to many systems compared to alternatives like Claude.
  • Apple Ecosystem Analogy: The presenter compares this limitation to the Apple ecosystem, where users are locked into specific services.

4. Writing Capabilities

  • Social Post Example: GPT5 can write social media posts, but without specific instructions, it may produce content with noticeable "AISMs."
  • M-Dash Usage: The model still uses M-dashes excessively.
  • Improvement Over GPT-4: GPT5 is better than GPT-4 for writing, but still requires fine-tuning for optimal results.

5. API Pricing and Model Comparison

  • Affordable API: The GPT5 API is priced at the same level as Gemini 2.5 Pro but offers a larger context window (400,000 tokens).
  • GPT5 Mini: The GPT5 Mini model is very cheap and only slightly less smart than GPT5 (around 20% difference).
  • Price Comparison: GPT5 Mini is five times less expensive than GPT5 for both input and output.
  • Competitive Pricing: GPT5 Mini is cheaper than Google's 2.5 Flash and significantly cheaper than Entropic's Opus 4.1.

6. Real-World Tests and Results

  • Test 1: Difficult Situation Email:
    • Scenario: Responding to a frustrated client about a delayed marketing automation system.
    • Models Tested: Gemini 2.5 Flash, GPT5, GPT5 Mini, Nano, Opus, Sonnet, Gemini 2.5 Pro, DeepSe.
    • Winner: GPT5 produced the best email with a professional tone and actionable analysis.
    • Key Quote: "Provided the information is correct, I could almost send this email as is."
  • Test 2: Ad Analysis:
    • Scenario: Analyzing Meta Ads data to identify winning strategies.
    • Models Tested: Gemini 2.5 Pro, GPT5.
    • Winner: GPT5 provided clear, data-backed insights, while Gemini 2.5 Pro was overly verbose and lacked specific numbers.
  • Test 3: Vibe Coding (Retirement Planner):
    • Scenario: Creating a retirement planner with Monte Carlo simulations.
    • Models Tested: Gemini 2.5 Pro, Opus, GPT5.
    • Winner: GPT5 produced a visually appealing and functional calculator, surpassing Gemini 2.5 Pro and rivaling Opus.
  • Test 4: Product Launch Plan:
    • Scenario: Planning a product launch for a cohort program.
    • Models Tested: GPT5, Claude, Opus, Sonnet, Gemini 2.5 Pro.
    • Winner: GPT5 provided the most specific and actionable plan, graded as the best by GPT5 itself.
  • Test 5: LinkedIn Post:
    • Scenario: Writing a LinkedIn post to brag about client success.
    • Models Tested: GPT5, GPT5 Mini, Opus, Sonnet, Gemini 2.5 Pro, DeepSe.
    • Winner: GPT5 and Gemini 2.5 Pro were competitive, with GPT5 offering a more direct approach.
  • Test 6: Meeting Transcript Extraction:
    • Scenario: Extracting important data from a meeting transcript.
    • Models Tested: GPT5 Mini, Gemini 2.5 Flash.
    • Winner: GPT5 Mini was less verbose, provided more action items, and wrote better follow-up emails than Gemini 2.5 Flash.

7. Conclusion

  • Excellent Update: GPT5 is an excellent update from OpenAI, benefiting both casual users and API developers.
  • Affordable and Powerful: The models are powerful and affordable, especially GPT5 Mini.
  • GPT5 Mini as Workhorse: GPT5 Mini is ideal for automation and API tasks due to its cost-effectiveness.
  • Worth Trying: Even if you use other chatbots, GPT5 is worth trying due to its capabilities and free availability.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.