I Tested Every New AI Model for Client Work - Here's the Winner

Arseny ShatokhinAbout 6 min readAug 19, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • LLMs (Large Language Models): GPT-5, Gemini 2.5 Pro/Flash, Claude Opus/Sonnet, Grok
  • AI Agents: DB Query Agent, Deep Research Agent, Newsletter Generation Agent
  • Real-world Use Cases: Analytics, Research, Content Generation
  • Light LLM: Framework for switching between different LLMs
  • Tool Calls: Interactions with external tools (e.g., web search, database query)
  • Benchmarks vs. Real-world Performance: Discrepancy between benchmark scores and actual task performance
  • Scoring Agent: AI agent used to evaluate the performance of other agents
  • Token Cost: Cost associated with the number of tokens used by an LLM
  • Latency: Time taken by an LLM to generate a response
  • Hallucinations: LLMs generating incorrect or nonsensical information
  • RAG (Retrieval-Augmented Generation): Improving LLM accuracy by grounding it with external knowledge

DB Query Agent (Analytics Use Case)

  • Description: An agent designed to query databases, particularly those with hundreds of unlabeled tables and unclear descriptions.
  • Structure: Single DB query agent with three tools.
  • Query Example: "What is the company's gross revenue from specific date to specific date?"
  • Evaluation: Scoring agent assesses alignment with criteria and accuracy of results.
  • Results:
    • GPT-5: Best performance in data analysis, but with high latency.
    • GPT-5 Mini: Similar performance to GPT-5, but uses double the tokens.
    • Gemini 2.5 Flash: Uses significantly fewer tokens (8x less than GPT-5 Mini).
    • Claude Opus: High cost, but not significantly better LLM score.
    • Grok: Did not stand out in any metric.
  • Overall Grade: Gemini 2.5 Pro ranked best (40% performance, 30% speed, 30% cost).
  • Conclusion: Gemini 2.5 Pro is best for analytics due to decent performance, high speed, and cost-effectiveness. GPT-5 excels in data analysis but suffers from high latency.

Deep Research Agent

  • Description: An agent that performs research on local files and the web.
  • Tools: MCP server and a web search tool (native to each model provider).
  • Research Query: "Compare Langraph vs. OpenAI Agents SDK vs. CrewAI vs. Pientic AI for production AI agents."
  • Files: Documentation for each framework.
  • Results:
    • GPT-5: Highest token usage (over a million), best performance (9/10).
    • GPT-5 Mini: Token usage is two times higher than Claude Sonnet 4.
    • Claude Opus: Failed badly.
    • Gemini 2.5 Flash: Low token usage (30,000), decent results.
  • Overall Grade: Gemini 2.5 Flash is the winner.
  • Conclusion: Gemini 2.5 Flash is best for research due to its efficient use of tokens and decent performance, leveraging Google's search capabilities. GPT-5 has the best performance but is very expensive.

Newsletter Generation Agent

  • Description: An agent that converts Figma files (website images and assets) into HTML. Based on a reverse-engineered Claude Code agent.
  • Process: Agent receives an image of a website and assets, then generates HTML.
  • Results:
    • Claude Opus: High cost ($7.50 per newsletter), webpage not well-aligned.
    • Claude Sonnet: Better alignment than Opus, website footer looks good.
    • Gemini 2.5 Flash: Misaligned images.
    • Gemini 2.5 Pro: Decent, but not as good as Claude Sonnet. Long completion time (over 10 minutes).
    • GPT-5: Almost perfectly matches the reference, excellent at UI generation. Social media icons are different sizes.
    • GPT-5 Mini: Also looks good, but social media icons are also screwed up.
    • Grok: Failed miserably.
  • Conclusion: GPT-5 excels at UI generation (front-ends). Claude Sonnet is preferred for back-ends.

Overall Impression of GPT-5

  • Assessment: GPT-5 is not as impressive as it would have been 6-9 months ago. Other models have caught up or surpassed it.
  • Cost: Relatively cheap, but high latency offsets the cost advantage.
  • Hallucinations: Significantly reduced compared to previous models.
  • Instruction Following: Excellent at following instructions precisely.
  • Coding: On par with Claude Sonnet, but slower.
  • Recommendation:
    • GPT-5: Analytics and RAG use cases, situations where hallucinations need to be minimized, and when precise instruction following is required.
    • Claude Sonnet: Coding use cases.
    • Gemini 2.5 Pro: General agentic use cases.
    • Gemini 2.5 Flash: Web browsing use cases.
  • Speculation: OpenAI may be cost-cutting due to a large number of free users. The current GPT-5 may not be the same as the one previously available to some users.

Significant Statements

  • "So, does GPT5 actually suck?" - Question posed at the beginning of the video, setting the stage for the comparison.
  • "Benchmarks don't really reflect the performance on real world tasks." - Highlighting the importance of real-world testing.
  • "GPT5 is exceptional at analyzing the data. However, the latency of this model is absolutely horrible." - Summarizing GPT-5's strengths and weaknesses in the DB Query Agent use case.
  • "One thing I did note about GPT5, however, is that the hallucinations for this model are significantly reduced." - Pointing out a key advantage of GPT-5.
  • "GPT5 always does exactly what you tell it to do." - Describing GPT-5's precise instruction following.

Technical Terms and Concepts

  • LLM (Large Language Model): A type of AI model trained on a massive amount of text data, capable of generating human-like text, translating languages, and answering questions.
  • AI Agent: An autonomous entity that perceives its environment through sensors and acts upon that environment through actuators.
  • Token: A basic unit of text used by LLMs for processing.
  • Latency: The delay between a user's input and the system's response.
  • Hallucination: A phenomenon where an AI model generates incorrect or nonsensical information that is not grounded in reality.
  • RAG (Retrieval-Augmented Generation): A technique for improving the accuracy and reliability of LLMs by grounding them with external knowledge retrieved from a database or the web.

Logical Connections

The video logically progresses through three distinct real-world AI agent use cases: DB Query Agent, Deep Research Agent, and Newsletter Generation Agent. For each use case, the video outlines the agent's purpose, structure, and the specific task it performs. It then presents a side-by-side comparison of various LLMs (GPT-5, Gemini 2.5 Pro/Flash, Claude Opus/Sonnet, Grok) based on performance, speed, and cost-effectiveness. The results are visualized through charts and reports, leading to an overall grade and recommendation for each use case. Finally, the video synthesizes the findings to provide an overall assessment of GPT-5's strengths and weaknesses, along with recommendations for when to use each LLM.

Data and Statistics

  • DB Query Agent: Cost comparison showing Claude Opus spending almost $3, significantly higher than other models.
  • Deep Research Agent: GPT-5 spending over a million tokens, while Gemini 2.5 Flash spent only 30,000 tokens.
  • Newsletter Generation Agent: Claude Opus spending $7 for a single newsletter.
  • OpenAI has 700 million weekly active users on Chat GPT, but only 3% of them are paid.

Synthesis/Conclusion

The video concludes that while GPT-5 exhibits strengths in data analysis, UI generation, and reduced hallucinations, its high latency and the emergence of comparable or superior models from other providers (particularly Gemini 2.5 Pro/Flash and Claude Sonnet) make it less of a game-changer than it would have been months ago. The video emphasizes the importance of considering specific use case requirements (analytics, research, coding, web browsing) when selecting an LLM, and provides tailored recommendations based on performance, speed, and cost-effectiveness. The speaker suggests that OpenAI may be prioritizing cost-cutting with GPT-5, and that the current version may differ from earlier iterations.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.