Sonoma (Grok 4 Mini) & Qwen-3 Max: 2 New Models are here but are they GOOD?

AICodeKingAbout 4 min readSep 7, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Soma Dusk & Soma Sky: Two new stealth models, likely from XAI (possibly Grok 4 Mini).
  • Quen Max: A new, non-open-sourced model available via API.
  • Context Window: The amount of text a model can consider at once (Soma models: 2 million tokens).
  • Tool Calling: A model's ability to use external tools or APIs.
  • Leaderboard: A ranking system used by the speaker to evaluate AI models.
  • Kilo Code: A platform where users can access and test AI models.
  • Ninja Chat: An all-in-one AI platform offering access to various top AI models.
  • GLM Models: A family of AI models that the speaker currently favors.

Soma Dusk and Soma Sky

  • Overview: Two new stealth models named Soma Dusk and Soma Sky. The speaker believes they are likely from XAI, potentially related to Grok 4 Mini.
  • Performance: Sky is considered the better model of the two, while Dusk is weaker. Both models exhibit some reasoning capabilities and support a 2 million token context window.
  • Tool Calling: Both models are decent at tool calling, further supporting the theory that they are related to Grok.
  • Leaderboard Ranking: Sky ranks 10th and Dusk ranks 15th on the speaker's leaderboard.
  • Specific Tasks:
    • Floor plan generation: Works, but the output is not impressive.
    • Panda SVG generation: Mediocre.
    • Autoplay chess game: Doesn't work properly; makes illegal moves.
  • Coding: Good at coding and tool calling. Can be integrated into Kilo Code.
  • Codebase Understanding: The 2 million token context window makes them potentially useful for codebase understanding.
  • Agentic Question Example: When asked to create a movie tracker app, the model immediately requested an API key for TMDB instead of using a placeholder, which the speaker found humorous. The initial code generation was followed by rate limit issues.
  • Recommendation: The speaker wouldn't personally use them for coding, suggesting GLM or DeepSeek as better alternatives. However, the large context window could be beneficial for code understanding.
  • Attribution: The models identify as "Grok from XAI" when prompted.

Quen Max

  • Overview: A new, non-open-sourced model available via API.
  • Availability: Can be used for free on Kilo Code with the $25 free credit.
  • Performance: The speaker is disappointed with Quen Max.
  • Leaderboard Ranking: Ranks 12th on the speaker's leaderboard, which is considered poor for a model that is not open-source, has over a trillion parameters, and costs more than GPT-5 Mini.
  • Tool Calling: Lacks in tool calling and doesn't work well.
  • Specific Tasks:
    • Floor plan generation: Not good.
    • SVG outputs: Not good.
  • Reasoning Issues: Attempts reasoning even though it's not a reasoning model, leading to crashes in tasks like mathematics questions (exceeds output limit).
  • Comparison to Previous Version: Quen 2.5 Max was considered "kind of cool," but this version is "pretty bad."
  • Comparison to Kimmy K2: Kimmy K2 is considered better due to its superior tool calling and fewer bugs.
  • Recommendation: The speaker cannot recommend Quen Max.

Ninja Chat Advertisement

  • Overview: Ninja Chat is presented as an all-in-one AI platform.
  • Features: Access to top AI models like GPT-4o, Claude 4 Sonnet, and Gemini 2.5 Pro.
  • Use Cases: Gemini is used for quick research. The AI playground allows side-by-side comparison of responses from different models. The mind map generator is useful for organizing complex ideas.
  • Pricing: $11 per month for the basic plan, which includes 1,000 messages, 30 images, and 5 videos monthly. Higher tiers are available.
  • Discount Codes: "king25" for 25% off any plan or "king40yearly" for 40% off annual subscriptions.

General Observations and Future Outlook

  • "Battle of Mid Models": The speaker characterizes the current state as a "battle of mid models," implying that none of the new models are truly exceptional.
  • Underrated Models: Models like Longat are considered underrated but lack provider support.
  • GLM Models: The speaker expresses a preference for GLM models due to their performance and the new coding plan.
  • Excitement for Gemini 3: The speaker is anticipating the release of Gemini 3 models, having already been impressed with Gemini 2.5 Pro.
  • Call to Action: The speaker encourages viewers to try the models and share their thoughts in the comments.

Synthesis/Conclusion

The video reviews three new AI models: Soma Dusk, Soma Sky, and Quen Max. The speaker is generally unimpressed with all three, considering them to be "mid models." Soma Sky is slightly better than Soma Dusk, and both are likely related to Grok 4 Mini from XAI. Quen Max is particularly disappointing, given its cost and parameter count. The speaker recommends exploring GLM models and expresses anticipation for the release of Gemini 3. The video also includes a promotion for Ninja Chat, an all-in-one AI platform. The main takeaway is that while new models are constantly being released, none of these particular models are game-changing.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.