Kimi K2 Reasoning (Fully Tested): IS IT REALLY THE BEST Open Model?

By AICodeKing

Share:

Key Concepts

  • Kim K2 Thinking: A new reasoning variant of the Kim K2 model developed by Moonshot AI, designed for complex tasks and step-by-step reasoning.
  • Agentic Benchmarks: Tests designed to evaluate a model's ability to perform tasks autonomously, using tools and planning over multiple steps.
  • Non-Agentic Benchmarks: Tests that assess a model's performance on individual tasks without requiring complex multi-step planning or tool usage.
  • Tool Calls: The ability of an AI model to interact with external tools or functions to gather information or perform actions.
  • State-of-the-Art (SOTA): Achieving the highest level of performance currently possible on a given benchmark or task.
  • Humanity's Last Exam (Browse Comp): A benchmark designed to test advanced reasoning and problem-solving capabilities.
  • GPT-5 CodeX: A model previously used by the speaker for planning and debugging tasks.
  • One Trillion Parameter Model: Refers to the large size and complexity of the Kim K2 Thinking model.
  • API Pricing: The cost associated with using the model's functionalities through an application programming interface.

Kim K2 Thinking: A New Reasoning Variant

Moonshot AI has introduced a new reasoning variant of their Kim K2 model, named "Kim K2 Thinking." This model is designed as a "thinking agent" capable of step-by-step reasoning and utilizing tools to achieve state-of-the-art performance on various benchmarks, including "humanity's last exam" (Browse Comp). The developers claim it exhibits major gains in reasoning, agentic search, coding, writing, and general capabilities.

Key Capabilities of Kim K2 Thinking

  • Sequential Tool Calls: Kim K2 Thinking can execute between 200 to 300 sequential tool calls without human intervention, demonstrating its ability to reason coherently across hundreds of steps to solve complex problems.
  • Planning, Reasoning, Execution, and Adaptation: While actively using a diverse set of tools, the model can plan, reason, execute, and adapt across hundreds of steps to tackle challenging academic and analytical problems.
  • Complex Problem Solving: An instance highlighted involved the model successfully solving a PhD-level mathematics problem through 23 interleaved reasoning and tool calls, showcasing its capacity for deep, structured reasoning and long-horizon problem-solving.
  • Performance Against Closed-Source Alternatives: The model reportedly achieves state-of-the-art performance on many benchmarks, even surpassing closed-source alternatives.

Benchmark Testing: Non-Agentic Tasks

The speaker conducted benchmarks on both non-agentic and agentic tasks. For non-agentic tests, the following results were observed:

  • Floor Plan Generation: The model failed to generate a floor plan, resulting in a blank white screen. Retries also led to errors.
  • SVG Panda with Burger: The generated SVG was described as "kind of bad" and not meeting expectations.
  • Pokeball in 3JS: This generation was considered "fine" and a "solid generation," although a black line incorrectly passed through the button.
  • Chess Game: The model performed well, consistently generating legal moves. While the UI was not optimal, the functionality was present, marking it as a "pass."
  • Minecraft in Kandinsky Style: This task was executed "really well." Minor issues included incorrectly placed trees and the absence of jumping mechanics, but it was deemed "solid."
  • Butterfly Flying in Garden Simulation: The simulation was "pretty good" and "really solid." The speaker desired more natural elements but found the generation "pretty bland."
  • CLI Tool in Rust: The generation worked "kind of well."
  • Blender Script: The script produced had incorrect syntax, resulting in a "fail."
  • Math Questions: Both math questions were unsuccessful, leading to a "fail."
  • Riddle: The model successfully solved a simple riddle, marking it as a "pass."

Overall Non-Agentic Performance: Kim K2 Thinking scored 13th position in the non-agentic benchmarks. It was noted to be slightly above Minimax but not as proficient in plain coding as Minimax, which is described as "quite better and faster."

Potential as a Planning Model

Despite not being the best in raw horsepower, Kim K2 Thinking is considered a strong contender to replace the speaker's current planning model (implied to be GPT-5 CodeX). As a one trillion parameter model with extensive knowledge of planning, it presents a significant alternative.

Benchmark Testing: Agentic Tasks

The speaker used Claude Code for agentic tests due to bugs in the Kimmy CLI with the new model, which would error out after 10-15 turns. Claude Code's implementation of interleaved thinking was utilized.

  • Movie Tracker App: The implementation was "very buggy." Navigation between pages resulted in errors, and attempts to fix them led to persistent issues. The speaker did not provide extensive feedback to maintain a one-shot benchmark approach. This was rated as "not great and not very usable."
  • Godo FPS Shooter Game: This task also performed poorly. Initially, it did not work. After error correction, the step counter was fixed, but the life bar did not update correctly after jumping as requested.
  • Spelta: The generation resulted in "a bunch of syntax errors" and did not work.
  • Tari App: Similar to Spelta, this also resulted in a "fail."
  • Go TUI Calculator: This was a success. The calculator was well-designed, correctly aligned, and functional.
  • Open Code Repo (SVG Generation Command): The task of adding an SVG generation command to an open code repository was a "fail."

Overall Agentic Performance: Kim K2 Thinking secured 10th position on the agentic leaderboard. Its performance footprint was noted as similar to GPT-5 CodeX.

Comparison with GPT-5 CodeX and Other Models

The speaker frequently uses GPT-5 CodeX for planning and considers it the best model for debugging and planning. Initial testing suggests that Kim K2 Thinking could potentially replace GPT-5 CodeX for planning tasks.

  • Debugging and Error Understanding: Larger models like Kim K2 Thinking are generally good at debugging due to their training on understanding and fixing numerous errors. They are also typically quick.
  • Writing Capabilities: Kim K2 Thinking is also noted as being great at writing, likely for similar reasons related to its extensive training data.
  • Chat Model Potential: The speaker finds Kim K2 models pleasant to interact with and believes Kim K2 Thinking could be used as a daily chat model, potentially replacing GPT-5 for this purpose as well.

Limitations and Recommendations

While Kim K2 Thinking shows promise, it is not recommended as the best coding model for daily use. The speaker intends to further test its capabilities in planning and report back.

API Pricing and Usability

The pricing for Kim K2 Thinking's API is as follows:

  • Slower API: $0.60 for input, $2.50 for output.
  • Turbo Variant: $1.15 for input, $8.00 for output.

The regular slower API is described as "almost unusable" due to extreme slowness. The turbo variant is "pretty snappy" and was used for testing. However, its high price makes recommending it difficult.

Conclusion and Future Outlook

Kim K2 Thinking is a "pretty cool" model with significant advancements in reasoning and tool usage. While it demonstrates potential as a planning and chat model, it falls short as a primary coding model. The high cost of the turbo variant is a deterrent for widespread recommendation. The speaker plans to continue evaluating its planning capabilities.

The video concludes with a call for viewer engagement, encouraging comments, subscriptions, and donations.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video