NEW Qwen 3, Better than Kimi K2?

Prompt EngineeringAbout 4 min readJul 23, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Openweight Models: AI models with publicly available weights.
  • Hybrid Reasoning Model: A single AI model capable of both thinking (reasoning) and non-thinking (direct response) modes.
  • Dedicated Reasoning/Non-Reasoning Models: Separate AI models, one optimized for reasoning tasks and the other for tasks requiring direct responses.
  • Benchmarks: Standardized tests used to evaluate the performance of AI models.
  • RKGI Score: A metric for evaluating the "real-world knowledge grounding" of a model.
  • Chain of Thought (CoT): An approach where the model explicitly shows its reasoning steps.
  • Tools/Code Interpreter: The ability of a model to execute code to solve problems.

Model Comparison: Kim K2 vs. Quent3 (Non-Reasoning)

The video compares the performance of Kim K2 (a leading openweight model) with Quent3, specifically the new 255B version with 22 billion active parameters, which is a dedicated non-reasoning model. The speaker also includes comparisons to proprietary models like Claude 4 Opus.

  • Quent3's Strengths:
    • State-of-the-art performance on several benchmarks when thinking is disabled.
    • Outperforms Kim K2 on some coding benchmarks (AD polyglot is close).
    • Impressive RKGI score (41.8), surpassing even Claude 4 Opus with thinking disabled (30%).
    • Excels in complex visual tasks like the bouncing heptagon ball simulation, maintaining 20 balls.
  • Kim K2's Strengths:
    • Website creation task: Kim K2 produced a more complete website with images, while Quent3 failed to load images.
    • Planet generation task: Kim K2 generated a planet with proper rotation and atmospheric shadows, while Quent3 had issues with landmass movement and shadow creation.

Case Studies and Examples

  • Website Creation: The prompt was to create a website with 25 legendary Pokémon, including their types, snippets, and images. Kim K2 succeeded, while Quent3 failed to load images.
  • Bouncing Ball Simulation: The prompt was to simulate 20 balls bouncing inside a spinning heptagon. Both models performed well, but Kim K2's output was considered more visually realistic.
  • Planet Generation: The prompt was to generate a planet with random geometry, biomes, textures, clouds, and other surface features. Kim K2 produced a more realistic and functional output with proper rotation and shadows.
  • Maze Solving: The prompt was to solve a maze by providing a step-by-step path. Both Kim K2 and Quent3 exhibited chain-of-thought behavior, but their proposed solutions were incorrect. Claude 4 Opus (with thinking enabled) came close to solving the maze. 03 solved the maze using code interpreter.

Reasoning Capabilities and Chain of Thought

  • Even though Quent3 is a non-reasoning model, it exhibited chain-of-thought behavior when attempting to solve the maze. The model showed internal monologue and reasoning steps, such as counting segments and analyzing connections between cells.
  • Kim K2 also displayed similar reasoning behavior, including backtracking and trying different paths.
  • The speaker notes that solving mazes effectively likely requires tools for interacting with the maze environment.

Claude 4 Opus Comparison

  • Claude 4 Opus (without thinking) produced less detailed output in the planet generation task compared to Quent3 and Kim K2.
  • Claude 4 Opus (with thinking enabled) adhered to the prompt more closely and produced a more responsive output.
  • Claude 4 Opus (with thinking enabled) came close to solving the maze, demonstrating its reasoning capabilities.

03 and Tool Usage

  • 03 successfully solved the maze by using a code interpreter and implementing a breadth-first search algorithm.
  • This highlights the importance of tools for complex problem-solving.

Notable Quotes

  • (Regarding Quent3's RKGI score): "Even if you disable thinking on cloud 4, it's only able to get to 30%, whereas this one gets to 41.8, which is pretty impressive."
  • (Regarding the maze-solving task): "But what was really surprising to me was that even clot for opus in the creos iteration was not able to solve this."
  • (Regarding 03's maze-solving ability): "Yeah it actually works which is pretty awesome right. So it can it actually shows the importance of tools that an agent have access to and it makes it a lot simpler compared to if you just want to do it in your head."

Conclusion

The video provides a detailed comparison of Quent3 (non-reasoning) with Kim K2 and Claude 4 Opus across various tasks. Quent3 demonstrates strong performance on benchmarks and specific visual tasks, while Kim K2 excels in tasks requiring image generation and planet creation. The maze-solving task highlights the limitations of non-reasoning models and the importance of tools for complex problem-solving. The speaker concludes that Quent3 is a decent model but refrains from definitively calling it state-of-the-art based on the limited testing. The speaker also expresses curiosity about why the developers chose to create dedicated reasoning and non-reasoning models instead of retraining a hybrid model.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.