Claude 4.5 Haiku (Fully Tested): The WORST Model Anthropic has ever made! Scores #34 on KingBench

By AICodeKing

Share:

Key Concepts

  • Claude 4.5 Haiku: Anthropic's new "small" AI model, positioned as a cheaper and faster alternative to previous state-of-the-art models.
  • Claude Sonnet 4: A previous state-of-the-art model from Anthropic, used as a benchmark for comparison.
  • Drop-in Replacement: A component or model that can be substituted for another without requiring significant changes to the system.
  • Agentic Tests: Tests designed to evaluate an AI model's ability to perform tasks autonomously or as part of a system.
  • Benchmarks: Standardized tests used to measure and compare the performance of AI models.
  • Enterprise Use: The application of AI models by businesses and organizations.
  • API (Application Programming Interface): A set of rules and protocols that allows different software applications to communicate with each other.
  • Input/Output Tokens: Units of text (words, sub-words, or characters) that AI models process.
  • Valuations: The estimated worth of a company, often influenced by investor perception and performance metrics.

Claude 4.5 Haiku: Launch and Claims

Anthropic has launched Claude 4.5 Haiku, described as their latest "small" model, now available to all users. The company claims that what was recently at the frontier of AI is now more affordable and faster. Specifically, they state that Claude Haiku 4.5 offers similar coding performance to Claude Sonnet 4 (a state-of-the-art model from five months prior) but at one-third the cost and more than twice the speed. It is presented as a drop-in replacement for existing models on Claude Code and their apps.

Performance Evaluation: A Critical Assessment

Despite Anthropic's claims, the video presents a starkly different assessment of Claude 4.5 Haiku's performance, particularly in coding and visual generation tasks.

1. Visual Generation and Layout Tasks:

  • Floor Plan Generation: The model produced a floor plan that was deemed nonsensical and structurally unsound, with walls misplaced from any logical perspective.
  • SVG Panda Holding a Burger: The generated SVG was described as "really bad," with the panda recognizable but the overall layout and composition being poor.
  • 3JS Pokeball: The 3JS rendition of a Pokeball was also considered "pretty bad" and not useful.
  • Chessboard: The generated chessboard was similarly evaluated as "really bad."
  • Web Version of Minecraft: The attempt to create a web version of Minecraft was also deemed "really bad."
  • Butterfly in the Garden: This was the only visual task that was considered "fine," though not extraordinary.

2. Code Generation and Functionality:

  • CLI Tool and Blender Script: These generated scripts did not even work, indicating a fundamental failure in code generation.
  • Agentic Tests (using Claude Code):
    • Movie Tracker App: The app failed to display, resulting in a 404 error, which is considered "insanely bad" for a drop-in replacement.
    • Go Terminal Calculator: This application was described as "atrocious," producing errors, having poor layout, and being generally bad.
    • Godo Game: The model failed to create a functional Godo game, showing errors throughout.
    • Open Code Repo: The generated code for an open code repository was deemed "pretty bad" and "super bad."
    • Spelt and Nury: Performance in these areas was also described as "super bad," contributing to its ranking as one of the worst-performing AI coding agents.

The presenter states that Claude 4.5 Haiku scored the lowest on their tests among new-generation models from leading AI labs. They emphasize that they ran these tests multiple times with consistent negative results.

Comparison with Alternatives

The presenter contrasts Claude 4.5 Haiku with other models, suggesting that "GPT 5 Mini" (a hypothetical or benchmarked model) is an "insanely better alternative" in these specific tests. Other recommended alternatives for cheaper and better performance include:

  • GLM4.6: Described as an "insanely better model" with a cost of approximately $0.5 to $1.75 per million input and output tokens.
  • GPT5 Mini: Positioned as a strong alternative.
  • Gro Code Fast: Another recommended option for efficient coding.

Critique of Anthropic's Model Strategy

The video offers a critical perspective on Anthropic's recent model development and strategy:

  • Past Success vs. Current Performance: The presenter suggests that Anthropic's previous success with "3.5 Sonnet" might have been a "fluke or a stroke of luck," implying a decline in their ability to develop strong coding models since then.
  • Sonnet Series Evolution: Sonnet 3.7 was seen as the same model with reasoning added, Sonnet 4 was not a major improvement, and Sonnet 4.5 is considered a downgrade in many areas.
  • Model Structure and Pricing Discrepancy: Anthropic's model lineup (Opus, Sonnet, Haiku) is compared to OpenAI's (GPT5, GPT5 Mini, GPT5 Nano). Sonnet is equated to GPT5, and Haiku to GPT5 Mini. The absence of a "Nano" equivalent is noted, as mini and nano models are typically used for tasks like structuring and summarization.
  • Enterprise Focus and Benchmark Inflation: The presenter speculates that Claude 4.5 Haiku is designed to drive companies to use Anthropic's APIs, rather than for consumer use. They suggest the model is specifically trained on benchmarks to inflate performance numbers for investors, who prioritize these metrics over real-world utility. This strategy is seen as an attempt to increase company valuations.
  • Cost vs. Performance: The high cost of Claude 4.5 Haiku is highlighted, especially when compared to its significantly lower performance and the more affordable, better-performing alternatives like GLM4.6. The presenter states Haiku is three times higher in price for nearly 200% lower performance.

Conclusion and Recommendations

The presenter concludes that Claude 4.5 Haiku is a "really bad model" and cannot be recommended for any use, especially given its cost and the availability of superior alternatives. They express disappointment, stating they never expected such poor performance from a leading AI lab.

While Anthropic is reportedly working on fixes, the current assessment is overwhelmingly negative. The presenter advises users to try out the model for themselves using the free credits offered on Kilo Code ($25 in free credits) to form their own opinions.

The video ends with a call for viewer engagement, asking for their experiences and thoughts on the model, and encouraging subscriptions and support for the channel.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video