7 Coding LLMs, 1 Prompt—Here’s What I Found

Prompt EngineeringAbout 4 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Large Language Models (LLMs)
  • Web search tool integration
  • Information synthesis
  • Dashboard creation
  • Benchmarking (e.g., SUE Bench, ADER LLM Leaderboard, MMLU, HumanEval, Math)
  • Hallucination
  • Sequential tool call/Chain of Thought
  • Context window
  • Pricing per million tokens (input/output)
  • Agentic tasks
  • Model performance comparison

Model Testing and Setup

The video details a test of seven different LLMs using the same prompt to create a web app that synthesizes information from web searches into a dashboard. The goal is to assess their ability to follow instructions, use tools effectively, and avoid hallucinating information. The models tested include:

  • Claude Sonnet 4
  • Claude 3.7 (previous version)
  • Gemini 2.5 Pro
  • Quen 2.5 Max (with web search enabled)
  • DeepSeek R1
  • Groq 3 (dropped due to inability to function with the prompt)
  • Claude Opus 4

All models tested have web search capabilities and are reasoning models. A key differentiator is the sequential tool call/chain of thought capability, particularly in Claude Opus 4 and 03, allowing them to make multiple web search calls and build information iteratively.

Pricing Comparison

The video highlights the pricing differences among the models, which significantly impacts their practicality:

  • Opus 4: $75 per million output tokens
  • Claude Sonnet: $15 per million output tokens
  • 03: $40 per million output tokens
  • Gemini 2.5 Pro: $15 per million output tokens (best pricing, even better for <200k tokens)

Gemini 2.5 Pro emerges as the most cost-efficient option.

Results Analysis: Model-by-Model Breakdown

The video presents the results of each model, numbered 1 through 7, without initially revealing which model produced which output. The analysis focuses on:

  • Accuracy of Information: Whether the model correctly identifies state-of-the-art models, release dates, and model sizes.
  • Tool Usage: How effectively the model uses the web search tool and synthesizes information.
  • Hallucinations: Instances where the model fabricates information or provides incorrect benchmark scores.
  • UI Rendering: Whether the model correctly renders the web page and its elements.

Model 1 (Gemini 2.5 Pro):

  • Identifies Opus and Claude 4 series but misses 03.
  • Incorrectly identifies Jamba 1.5 Large as state-of-the-art.
  • Release dates are off by one day (likely due to timezone issues).
  • Benchmarks are presented with a visual component.

Model 2 (Sonnet 4):

  • Lists only Claude 4, not Opus 4.
  • Incorrectly lists an 8 billion parameter model as a frontier model.
  • Incorrect release date for Claude 4.
  • Incorrectly states Claude 4 has 200 billion parameters and Gemini 2.5 Pro has 1 trillion.
  • Benchmark tab is non-functional.

Model 3 (Opus 4):

  • Lists only Claude 3.7, not Claude 4.
  • Correctly lists 03.
  • Correct context window information.
  • Incorrectly identifies Falcon 2 as a frontier model.
  • Benchmarks are presented with plots, but specific benchmarks cannot be selected.
  • Lacks MMLU, HumanEval, and Math benchmarks for 03.
  • Allows filtering by company.

Model 4 (Sonnet 3.7):

  • Lists GPT 4.5 (non-existent), 03 Mini, and Claude 4.
  • Hallucinates benchmark scores (e.g., GPT 4.5 MMLU score).
  • Benchmark selection is functional, but some links are broken.

Model 5 (Quen 2.5 Max):

  • Limited information provided.
  • Incorrectly states Gemini 2.5 Pro has 540 billion parameters and a 16,000 token context window.
  • Lists benchmarks for the models provided in the prompt.

Model 6 (03):

  • Finds Claude Opus 4 but misses 03.
  • Lists DBRX as state-of-the-art (released in 2024).
  • Correct release dates.
  • Functional benchmark selection.
  • Visually unappealing.

Model 7 (DeepSeek R1):

  • Failed to render correctly and was discarded.

Sequential Tool Call Example (Claude Opus 4)

Claude Opus 4 demonstrates sequential tool call by:

  1. Performing a web search.
  2. Analyzing the results.
  3. Identifying different themes.
  4. Conducting subsequent web searches based on the identified themes.

This iterative process allows Opus 4 to build up its information gradually.

Conclusion

Despite providing the same prompt, the results across all models were inconsistent. All models except DeepSeek R1 rendered the UI correctly. Information synthesis was lacking even in Opus 4 and Gemini 2.5 Pro.

The video suggests that for complex tasks involving multiple agents, it's better to use a combination of models rather than relying on a single model. The presenter is biased towards Gemini 2.5 Pro due to its cost-efficiency, but performance-wise, Sonnet models or Opus could be considered, keeping in mind the potential for rate limits. The key takeaway is that irrespective of the model used, results should be rechecked, especially in multi-agentic systems.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.