Qwen 3.5 The GREATEST Opensource AI Model That Beats Opus 4.5 and Gemini 3? (Fully Tested)

WorldofAIAbout 5 min readFeb 17, 2026Watch original
THE SUMMARYAI-generated

Quen 3.5 Series: A Detailed Overview

Key Concepts:

  • Quen 3.5: Alibaba’s new flagship open-source multimodal AI model.
  • Openweight Model: A model with publicly available weights, enabling community access and modification.
  • Multimodal: Capable of processing and generating multiple data types (text, image, video, etc.).
  • Hybrid Linear Attention: A novel attention mechanism combining efficiency with performance.
  • Sparse Mixture of Experts (SMoE): An architecture that selectively activates parts of the model for different inputs, improving scalability.
  • Reinforcement Learning from Human Feedback (RLHF): A technique used to align the model’s behavior with human preferences.
  • Apache 2.0 License: A permissive open-source license allowing broad usage and modification.
  • Context Window: The amount of text a model can consider at once (Quen 3.5 supports 1 million tokens).
  • MMLU Pro & Video MME: Benchmarks used to evaluate model performance on massive multitask language understanding and multimodal tasks respectively.
  • Sway & Terminal Benchmarks: Benchmarks used to evaluate code generation capabilities.

1. Introduction & Model Overview

The Alibaba team has released the Quen 3.5 series, their first flagship openweight model. This is a 397 billion parameter model with 17 billion active parameters, designed as an all-purpose, natively multimodal system. It leverages a new architecture combining hybrid linear attention with sparse mixture of experts (SMoE) and is scaled using large reinforcement learning environments. A key design goal is its suitability for real-world agents. The model is reportedly 19 times faster than the previous Quen 3 Max and supports 201 languages and dialects.

2. Performance Benchmarks & Comparisons

Quen 3.5 demonstrates strong performance across various benchmarks:

  • MMLU Pro: 87.88
  • Video MME: 87.5
  • Browser Comp: Outperforms Claude Opus 4.5.
  • Multimodal Tasks: Surpasses Gemini 3 Pro in several areas.
  • AA Coding (Sway Bench): Matches Opus, exceeds Gemini 3 Pro.
  • AA Coding (Terminal Bench): Trails behind Gemini 3 Pro.

These results position Quen 3.5 as a leading openweight multimodal model.

3. Licensing & Accessibility

Quen 3.5 is released under the Apache 2.0 license, which is highly favorable for open-source development and multimodal agent capabilities. Access is available through multiple avenues:

  • Free Weights: All model sizes are publicly available.
  • Chatbot: A direct interface for interacting with the model.
  • API: Accessible via Alibaba’s cloud service.
  • Kilo Code: Offers a free API with up to $25 in credits.
  • Open Router: Another platform for accessing the model via API.

4. Strengths & Weaknesses

Strengths:

  • Speed & Efficiency: Significantly faster than Quen 3 Max.
  • Large Context Window: Supports a 1 million token context.
  • Multimodal Capabilities: Excels in vision plus reasoning and tool use.
  • Open Source: Facilitates community contributions and customization.

Weaknesses:

  • Complex Spatial Tasks: Struggles with intricate spatial reasoning.
  • Code Generation Consistency: Code generation is decent but not always reliable.
  • Real-World Stability: Less stable in real-world applications compared to top closed-source models.

5. Demonstrations & Use Cases

Several demonstrations showcased Quen 3.5’s capabilities:

  • 3D Mapbox Generation: The model autonomously generated a live, interactive 3D map of Beijing and Shanghai using React, including UI components and production-ready front-end logic.
  • Game Development (Super Mario Platformer): The Quen 3.5 Next Coder Q8 (8 billion parameters) successfully generated a functional Super Mario platformer game.
  • Game Development (Car Racing Game): The model generated a visually impressive car racing game, though the origin of the demo is unclear (potentially heavily prompted/edited).
  • Mac OS Browser OS: Generated a visually similar Mac OS interface in 20 cents using Kilo Code, initially with limited functionality, but improved with re-prompting.
  • SVG Generation: Successfully generated both simple and photorealistic animated butterflies.
  • Landing Page Creation: Generated a front-end landing page faster than Opus 4.6, demonstrating speed in generating typography, animations, and dynamic elements.
  • Multimodal Analysis: Accurately counted 28 toy cars in an image, demonstrating reasoning capabilities.
  • Farming Simulation Game: Generated a functional farming simulation game resembling Stardew Valley, including harvesting, planting, and animal interactions.
  • 3D Room Designer: Created a 3D room designer tool allowing furniture placement and color changes, though movement of objects was limited.
  • Video Generation: Generated a short video clip of two individuals engaging in a conversation.

6. Observations on Training & Development

The speaker noted a trend among Chinese AI labs of training their models on the output of proprietary models like Gemini. While not inherently negative, this practice raises questions about the originality of the research and development. The speaker stated, “Now, in my opinion, what it does look like is that these guys trained off of the Gemini output, which is something that a lot of these Chinese companies have been doing. I'm not shit-talking or anything, but I'm just stating the facts.”

7. Sponsorship: Mammoth AI

The video included a sponsorship segment for Mammoth AI, a platform providing access to multiple leading AI models (GPT, Gemini, Deepseek, Grock, etc.) through a unified API. Mammoth AI offers flexible pricing and supports various applications like deep research, code generation, and multimodal workflows.

8. Conclusion

Quen 3.5 represents a significant advancement in openweight multimodal AI. While it has limitations in complex spatial tasks and real-world stability, its speed, efficiency, large context window, and open-source licensing make it a compelling option for developers and researchers. It is currently considered the best open-source multimodal model available and a valuable tool for vision-language tasks, tool use, and potentially real-world agent applications. The speaker concludes, “Overall, this is a great remarkable multimodal model and it's something that you can potentially use for realworld agents.”

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.