Kimi K2.1 (0905 - Fully Tested): Does this REALLY BEAT Claude 4!?

AICodeKingAbout 4 min readSep 5, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Kimmy 0905: An improved version of the Kimmy K2 model.
  • Context Window: The amount of information a model can consider at once (increased to 256k tokens).
  • Tool Calling: The ability of a model to use external tools or APIs.
  • Agentic Tasks: Complex tasks requiring multiple tool calls and reasoning.
  • Raw Benchmarks: Tests measuring the model's general intelligence and performance.
  • Ninja Chat: An all-in-one AI platform offering access to multiple AI models.
  • Temperature: A parameter controlling the randomness of the model's output (recommended 0.6 for Kimmy 0905).
  • GLM Coding Plan, Claude Code, Codeex: AI coding assistants used for comparison.
  • TMDB API: A movie database API.
  • Bubble Tea: A Go framework for building terminal-based UIs.
  • Godot: A game engine.

1. Introduction of Kimmy 0905

  • Kimmy has launched a new model, Kimmy 0905, an improved version of the Kimmy K2 model.
  • The presenter received early access to the model.
  • The model should be available for users to try out.

2. Pricing

  • The base model is priced at $1.50 and $2.50.
  • The turbo version (faster inference) is priced at $2.40 and $10.
  • The presenter primarily tested the base low-cost version.
  • The weights have been released, allowing price comparison across providers.

3. On-Paper Upgrades

  • Context Window: Increased to 256k tokens. This is a significant and useful upgrade.
  • Coding Skills: Improved, especially for front-end development and tool calling.
  • Tool Calling: Better at using tools like Claude Code, Ruode, and Codeex.
  • Raw Intelligence: Approximately a 10% improvement over the previous model.
  • Benchmarks: Generally a 10% improvement in most benchmarks.

4. Ninja Chat Advertisement

  • Ninja Chat is an all-in-one AI platform for $11 per month.
  • Access to top AI models like GPT40, Claude for Sonnet, and Gemini 2.5 Pro.
  • AI playground for comparing responses from different models.
  • Mind map generator for organizing complex ideas.
  • Basic plan: 1,000 messages, 30 images, and 5 videos monthly.
  • Discount codes: "king25" for 25% off any plan, "king40yearly" for 40% off annual subscriptions.

5. Raw Benchmarks Results

  • Kimmy 0905 ranks 10th in the presenter's raw benchmarks.
  • An improvement from the previous version.
  • Better than Sonnet, slightly worse than GLM.
  • Floor Plan Generation: "Kind of good," but walls don't make sense.
  • SVG Generation: "Not super great."
  • Chessboard: Functionality works, but doesn't make legal moves or check the king.
  • Butterfly: "Quite good," with animation.
  • CLI Tool for Image Conversion: "Great."
  • Blender Script: Doesn't work.

6. Agentic Tasks Performance

  • The model is better at agentic tasks with multiple tool callings.
  • Specifically improved in Claude Code, Ruode, and Clin.
  • Can be used for free via Kilo Code with $25 free credits.

7. Agentic Test Examples

  • Movie Tracker App (Expo, TMDB API):
    • Kimmy: "Did it kind of well." Movies look nice, inner page is "wonky," calendar is "pretty good."
    • GLM: "Looks pretty cool. Like really cool." More usable, fewer issues.
    • Claude: "Great and it works well."
    • Codeex: "Good but very lackluster." Incorrect colors, title bar left as is.
  • Go-Based Terminal Calculator (Bubble Tea):
    • Kimmy, Claude, Codeex: Fail.
    • GLM: Succeeds, builds a "good-look thing."
  • Godot Game Edit (Step Tracker):
    • Kimmy, Codeex, GLM: Fail.
    • Claude: Passes with ease in one shot.
  • SVG Generation Command (Open Code Repo):
    • All models fail.

8. Overall Assessment

  • Kimmy 0905 scores third position in agentic tests, above Codeex but below GLM and Claude Code.
  • Claude Code is the best in these tests.
  • A better option than the previous Kimmy K2 model.
  • One of the biggest open models with one trillion parameters.
  • Recommended temperature setting: 0.6 for best generations.
  • Reliable in tool calling.
  • Performance based on density is not as great as desired.
  • Hopes for improvement in Kimmy 3 or 2.5.
  • GLM is currently the best AI coder for the presenter.

9. Upcoming Video

  • A video comparing the GLM coding plan with Codeex and Claude Code is coming soon.

10. Conclusion

  • Kimmy 0905 is a good model, especially for tool calling, but its performance relative to its size could be better. It needs to catch up to smaller, more efficient models. The presenter anticipates further improvements in future iterations.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.