NEW Qwen 3 Coder: Did the Benchmark Lie?

Prompt EngineeringAbout 3 min readJul 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Quen team and their open-source/openweight models
  • Quen 3 series coding model
  • Model architecture: Mixture of Experts (MoE)
  • Context window size and its implications
  • Agentic coding, browser use, and tool use
  • Pre-training data and scaling
  • Reinforcement Learning (RL) post-training
  • Reasoning vs. non-reasoning models
  • Sweep Bench Verified benchmark
  • RKGI score and its verification
  • Model testing and examples: UI creation, animations, web apps

1. Introduction to Quen Coding Model

  • The Quen team has released a new coding model based on their Quen 3 series.
  • This model is significant because it's an openweight model that rivals Claude Sonnet 4 on key benchmarks.
  • Early testing suggests it's a solid coding model with unique capabilities.

2. The Quen Team and Their Contributions

  • The Quen team is a major contributor to open-source and openweight models, comparable to DeepSeek and Kimiku.
  • They have various models, including the Quen 3 series, re-rankers, vision language models, and Omni models (multimodal).
  • Nvidia's Neotron open reasoning family is built on top of Quen models, highlighting their strength.
  • Quen also open-sources Quen Code, built on Gemini CLI, similar to Cloud Code.

3. Model Details and Technical Specifications

  • The model is a 480 billion parameter model, but only 35 billion parameters are active (Mixture of Experts).
  • It has a context window of 256k tokens, extendable to 1 million tokens, comparable to Gemini series.
  • Specifically trained for coding, making it suitable for agentic coding, browser use, and tool use.

4. Training Data and Methodology

  • The model was trained on 7.5 trillion tokens, with about 70% being code.
  • Quen believes there's room to scale in pre-training, unlike Kim K2, which uses more tokens.
  • They use synthetic data generated by Quen 2.5 coder for cleaner pre-training data.
  • RL is used during post-training, with potential for further improvement.

5. Reasoning vs. Non-Reasoning Approach

  • The model is a non-reasoning model, similar to their updated Quen 3 model.
  • It achieves high performance on Sweep Bench Verified, comparable to Claude Sonnet 4, without test-time scaling or reasoning.
  • They created 20,000 independent environments in parallel on Alibaba Cloud for long-horizon RL training.
  • ARC AGI research suggests that shorter thinking durations can lead to more accurate answers, questioning the necessity of long reasoning chains.

6. Benchmark Scores and Verification

  • The Quen team claimed a 41.8% RKGI score for the updated Quen 3 release, which was questioned by Francis Shraw from the RKGI team.
  • The RKGI team couldn't reproduce the claimed score on public or semi-private sets.
  • The Quen team is assisting in reproducing the RKGI score.
  • Benchmarks should be taken with a grain of salt and tested on private datasets.

7. Practical Examples and Testing

  • The model excels at UI creation and instruction following as a coding agent.
  • Examples include:
    • Balls falling within a spinning heptagon with added controls and animations.
    • A web app for a gallery of visited places with simulated data.
  • The model struggles with solving mazes but could perform better in an agentic mode with coding tools.

8. Availability and Usage

  • The model is available on the Quen platform, Hugging Face, Open Router, and AnyCoder.
  • It can be used within Quen Code or Cloud Code.
  • Open Router provides a free API key for use within clients like Client.

9. Conclusion

  • The Quen coding model is an impressive release, especially for UI creation and animation tasks.
  • Its performance on benchmarks is notable, but verification is crucial.
  • The model's non-reasoning approach and focus on efficient training are interesting choices.
  • Testing the model and providing feedback is encouraged.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.