Key Concepts
- Quen 3 Coder: A 480B parameter coding model by Quen, excelling in agentic coding, browser use, and tool use.
- Mixture of Experts (MoE): An architecture where only a subset of the model's parameters (35B in this case) are activated for each input.
- Quen Code: A command-line tool for agentic coding, forked from Gemini Code and adapted for Quen 3 Coder.
- Agentic Coding: Coding tasks performed by an AI agent, often involving planning, tool use, and iteration.
- Context Window: The amount of text a model can consider when generating a response (256K or 1M tokens in this case).
- SWE-bench: A benchmark for evaluating coding models.
- ARC AGI: A benchmark for evaluating general AI capabilities.
- Open Router: A platform for accessing various AI models through a unified API.
- Klein, RU, Kilo: Development environments or platforms for using AI models.
Quen 3 Coder Model Overview
The video discusses the newly released Quen 3 Coder model, a 480B parameter model that stands out in agentic coding tasks. It's a mixture of experts model, activating only 35 billion parameters at a time, similar to models like Kimmy and Deepseek. While multiple sizes are planned, only the 480B model is currently available. The model excels in agentic coding, agentic browser use, and agentic tool use, showing performance comparable to Claude Sonnet 4 among open models.
Quen Code CLI Tool
Alongside the model, Quen has launched Quen Code, a command-line tool for agentic coding. It's based on Gemini Code but customized with specific prompts and function calling protocols to leverage the full capabilities of Quen 3 Coder.
Training Data and Benchmarking
The Quen 3 Coder model was trained on 7.5 trillion tokens. On the SWE-bench verified benchmark, it scores above Kimmy K2 but slightly below Sonnet 4. The speaker tested the model on five agentic tasks, comparing it to Claude Code, Gemini CLI, and Kimmy K2 with Claude Code router. In these tests, Claude Code and Kimmy K2 with Claude Code router solved three tasks each, while Gemini CLI and Quen Code solved two each. The speaker notes that Quen's primary limitation is its 256K context window.
Comparison with Kimmy and Other Models
The speaker finds Kimmy slightly better for their use cases, describing it as a local version of Sonnet. Kimmy tends to generate more verbose outputs, including reasoning in the output, which can increase costs. Quen, on the other hand, is more straightforward. The speaker emphasizes the need for more extensive use before making a definitive judgment but notes that both models are comparable to Sonnet. The advantage of open-weight models like Quen and Kimmy is the potential for faster speeds through platforms like Grok, contrasting with Anthropic's slower speeds.
Benchmark Concerns
The speaker raises concerns about Quen's benchmarking practices, referencing a callout from an ARC AGI benchmark author. Quen claimed a 41% score on ARC AGI for their 235B model, which the benchmark authors couldn't reproduce. The speaker notes that it's standard practice for benchmark authors to receive a private endpoint for testing and verifying results. The speaker expresses distrust in Quen's benchmark scores due to their alleged heavy training on benchmark questions, a practice that Kimmy and DeepS do not currently engage in.
Usage and Availability
The Quen 3 Coder model is available on the Quen chat platform for free, including a webdev option for creating React applications. It's also accessible through Open Router, with two versions: one with a 1 million token context window and another with a 256K context window.
Pricing and API Considerations
The speaker strongly advises against using Quen's official API due to its high cost. The API pricing ranges from $1 to $6 based on the total context used, with costs potentially doubling or tripling per token if the context window tier is exceeded. The output token cost for 1 million tokens input can reach $60, comparable to Opus, which the speaker deems unreasonable. The 256K context window version costs $22, which is also considered too high. The speaker recommends using third-party providers like Hyperbolic (around $2 for input and output) or Parasale (around $2 and $3.50) for significantly lower costs.
Integration with Development Environments
The video explains how to use Quen 3 Coder with development environments like Klein, RU, and Kilo through VS Code and Open Router. It also mentions Requesty as an alternative. Kilo Code offers $20 in free credits for trying the model.
Quen Code CLI Usage
The speaker provides instructions on installing and using the Quen Code CLI tool. It can be installed from their repository using the provided command. The CLI is similar to Gemini CLI in functionality. The speaker notes that sometimes there is an error in the interactive setup mode, so you can just export the OpenAI API key and base URL environment variable and then use it that way.
Conclusion
The Quen 3 Coder model is a powerful tool for agentic coding, comparable to models like Claude Sonnet and Gemini CLI. However, users should be cautious about the official API pricing and consider using third-party providers for cost-effectiveness. The Quen Code CLI provides a convenient way to leverage the model's capabilities. The speaker expresses hope for the release of smaller models in the future.
AI summaries can miss context or contain errors. Check important details against the original video.