Key Concepts
- Qwen 3 Series: New family of AI models launched by Qwen (Alibaba).
- Mixture of Experts (MoE): Model architecture where different parts ("experts") specialize in different tasks, only activating relevant experts for a given input, leading to potentially faster inference with fewer active parameters.
- Hybrid Thinking Models: Models that exhibit a "thinking" process or chain-of-thought reasoning before providing a final answer, often visible in their output traces.
- Active Parameters: In MoE models, the number of parameters actually used during inference for a specific input, typically much lower than the total parameter count.
- Benchmarks: Standardized tests used to evaluate AI model performance (e.g., Aider for coding, math, general capabilities).
- APIs (Application Programming Interfaces): Interfaces allowing software applications to interact with the models (e.g., via OpenRouter).
- Local Usage: Running models directly on user hardware using tools like Ollama or VLLM.
- Klein / Rukode: Software tools mentioned for interacting with LLMs, likely IDE extensions or similar coding assistants.
- OpenRouter: A platform providing access to various LLM APIs, including free tiers for some models.
- Ollama / VLLM: Frameworks/tools for running large language models locally.
- Photogenius AI: Sponsor product - an AI art generation platform.
Introduction to Qwen 3 Models
Qwen has released its Qwen 3 series, comprising six models. Two are Mixture of Experts (MoE) models, while the other four are standard dense models. The primary focus has been on the largest MoE model (235B parameters) and the 32B general-purpose model.
Model Specifications and Benchmarks
- Qwen 3 235B MoE:
- Total Parameters: 235 billion
- Active Parameters: Approximately 22 billion.
- Type: Mixture of Experts.
- Claimed Performance: Achieves competitive results against models like Deepseek R1, Sonnet 3 Mini, Grok 3, and Gemini 1.5 Pro in coding, math, and general capabilities benchmarks.
- Claimed Aider Score: ~61.8, potentially on par with Sonnet (in non-thinking tasks). The speaker expresses skepticism, doubting it surpasses Deepseek R1 given the size difference, and notes benchmark scores can be influenced by system prompts or specific training data (like the 32B coder potentially being trained heavily on benchmark questions).
- Local Usage: Unlikely without quantization due to size.
- Qwen 3 32B General:
- Total Parameters: 32 billion.
- Type: Standard dense model.
- Claimed Performance: Boasts comparable performance to the 235B model, though expectedly underperforms it.
- Qwen 3 30B MoE:
- Total Parameters: 30 billion.
- Active Parameters: Approximately 3 billion.
- Type: Mixture of Experts.
- Claimed Performance: Comparable to the Qwen-VL-Max model.
- Local Usage: Designed to be runnable locally, potentially offering fast inference due to the MoE architecture.
- Smaller Models:
- Sizes: 14B, 8B, 4B, 1.7B, 0.6B.
- Claimed Performance (4B): Apparently comparable to a 72B model, which the speaker finds "sketchy".
- General Assessment: Viewed as incremental improvements ("a bit better") over Llama 3.1 and previous Qwen versions.
Hybrid Thinking Feature
All models in the Qwen 3 series are described as "hybrid thinking models," similar to architectures like Deep Hermes. This means they often exhibit an explicit reasoning or "thinking" step before outputting the final response. This feature is supported in the VLLM framework for local execution, and the models available via OpenRouter APIs are also thinking-enabled.
Speaker's Testing and Evaluation
The speaker conducted tests using a custom-built tool (developed with Gemini 1.5 Pro).
- Performance Issues: The models exhibit the "same old Qwen thinking problem," where they get stuck in a thinking loop and eventually error out without providing a response.
- Comparison to Claims: The speaker finds the models are "not as good for sure" and "not comparable to Deepseek at all."
- Code Generation: Code quality is inconsistent; formatting can be off. Specific tests, like generating a hexagon using the 235B model, failed.
- Smaller Models: Considered only slightly better than previous generations and competitors like Llama 3.1.
- Overall Assessment: The models are not yet on the level of Deepseek. The speaker expresses a preference for the (currently closed-source) Qwen Max model, which they found more competitive.
Usage and Integration (APIs and Local)
- API Access (OpenRouter):
- Free APIs are available on OpenRouter for all Qwen 3 models (including 235B and 32B). Paid options also exist.
- Caveat: Using the API might not be ideal compared to alternatives like Gemini 1.5 Flash, especially given the Qwen models' slowness due to long thinking traces.
- Local Setup (Ollama/VLLM):
- Models can be run locally using tools like Ollama or VLLM.
- Process with Ollama:
- Install Ollama.
- Run
ollama run <model_name:size>(e.g.,ollama run qwen:32b) to download and load the desired model.
- Integration with Klein and Rukode:
- Step-by-step (API via OpenRouter):
- Ensure Klein/Rukode are updated.
- Obtain an OpenRouter API key.
- Navigate to Klein/Rukode settings.
- Select "OpenRouter" as the provider.
- Enter the API key.
- Choose the desired Qwen 3 model (speaker recommends the largest, 235B).
- Step-by-step (Local via Ollama/VLLM):
- In Klein/Rukode settings, select Ollama or VLLM as the provider.
- Configure the specific model to use.
- Note: The speaker warns that the models (especially via API) will have issues and can be slow.
- Step-by-step (API via OpenRouter):
Specific Model Assessments (Speaker's View)
- 235B MoE: Suffers from thinking loop errors and inconsistent performance.
- 30B MoE: Considered "cool" and potentially very fast for local inference due to low active parameters. Suitable for fine-tuning for specific tasks.
- 32B Model: Praised as "great at like basic coding" and potentially the "best local coder" currently available in the open-source space, despite not being state-of-the-art compared to closed models.
Sponsor Mention: Photogenius AI
The video includes a sponsorship segment for Photogenius AI:
- Description: An all-in-one AI platform for generating images, videos, and 3D models from text prompts.
- Features: Access to various generation models (Flux, Stable Diffusion, Google Imagen, V2 video, Cling), AI image editing tools (avatar generator, background removal, logo/emoji/thumbnail/icon generators).
- Pricing: Starts at $10/month.
- Discount: 25% off using the code
king25.
Conclusion/Synthesis
The Qwen 3 series represents an update to Qwen's model lineup, featuring MoE architectures and hybrid thinking capabilities. While benchmark claims position them competitively against top-tier models, the speaker's testing reveals practical issues like instability ("thinking problem") and performance inconsistencies, particularly when compared to strong competitors like Deepseek. The models are seen as improvements within the open-source landscape but don't bridge the gap to the best closed-source models. The 30B MoE model is highlighted for its potential local inference speed, and the 32B model is noted as a strong option for local coding tasks. Free APIs via OpenRouter and local deployment options make them accessible, but users should be aware of potential performance limitations and bugs. Overall, the speaker finds the release "pretty cool" but is not "super impressed," viewing them as intermediate models rather than state-of-the-art contenders like Deepseek. Fine-tuning the open-source models for specific tasks is suggested as a potential use case.
AI summaries can miss context or contain errors. Check important details against the original video.