What Makes Qwen 3 Max Thinking So Weird?

By Prompt Engineering

Share:

Key Concepts

  • Quinn 3 Max Thinking: A new version of Quinn's language model with enhanced reasoning capabilities.
  • Thinking Tokens/Budget: A configurable parameter in Quinn 3 Max Thinking that limits the number of tokens used for reasoning.
  • Interleaved Function/Tool Calls: The ability of the model to seamlessly integrate external tool usage (like web searches) within its reasoning process.
  • Misguided Attention: A phenomenon where LLMs fail to correctly interpret or apply a critical detail in a problem, often reverting to a standard or well-known version of the problem.
  • River Crossing Problem: A classic logic puzzle often used to test reasoning capabilities.
  • Heptagon Prompt: A complex coding challenge involving animation within a specific geometric shape, used to test programming and creative capabilities.
  • Agentic Coding Tools: AI tools that assist in the coding process.

Quinn 3 Max Thinking: An In-Depth Analysis

This video provides a detailed review of Quinn 3 Max Thinking, a new iteration of Quinn's language model that emphasizes reasoning capabilities. The reviewer expresses initial excitement due to a pre-release plot suggesting performance on par with models like GPT-5 and GPT-4, but finds the actual results to be surprisingly mixed.

Introduction to Quinn 3 Max Thinking

Quinn 3 Max Thinking has been released with limited official documentation. However, a plot shared in September for the earlier Max version indicated that the Thinking variant was expected to be highly competitive. The model is accessible on the Quinn chart, where users can enable "thinking" and adjust the "thinking tokens" or "thinking budget," with a maximum of 82,000 tokens available.

Strengths: Programming and Information Verification

1. Simple Game Creation: The model demonstrates proficiency in creating simple games. An example provided is "pixel pad," a drawing application with color selection and an eraser tool, showcasing control over a visual interface. The reviewer also tested a game creation prompt, which the model successfully executed with all requested controls, indicating decent basic programming capabilities.

2. Handling Conflicting Information and Web Search: Quinn 3 Max Thinking was tested on its ability to handle conflicting information and verify news. The prompt involved a fabricated announcement of "GPT-6" and associated rumors.

  • Model's Response: The model correctly identified the misconception about GPT-6, stating, "I need to clarify or correct a critical misconception. There's no GPT-6 announcement by OpenAI on this date."
  • Information Discovery: It then accurately reported on the release of "GPT OSS Safeguard," a family of open-weight models focused on safety, and provided suggestions for the article deadline, recommending verification of sources and contacting OpenAI's press team.
  • Comparison: The reviewer notes that the non-thinking version of Quinn 3 Max produced a similar, positive response, which initially fueled high expectations for the thinking version.

3. Advanced Reasoning Traces and Tool Integration: The "thinking traces" of Quinn 3 Max Thinking are described as concise compared to other models. A key feature highlighted is its ability to perform interleaved function or tool calls.

  • Process: The model can execute a web search, analyze the results, and then perform another web search or tool call to fill in gaps in its understanding.
  • Significance: This capability is deemed crucial for creating powerful agents, as it allows for dynamic information gathering and problem-solving within the reasoning process.

Weaknesses: Misguided Attention and Complex Coding Failures

1. Misguided Attention in Reasoning Problems: Despite its advanced reasoning capabilities, the model exhibits misguided attention.

  • Runaway Trolley Problem Modification: The reviewer presented a modified trolley problem where the five people on the main track were already dead.
    • Model's Correct Interpretation: Quinn 3 Max Thinking correctly identified the critical detail, stating, "Wait, that doesn't sound right... If they are already dead, the ethical dilemma collapses entirely." It correctly concluded not to pull the lever because the people were already deceased.
  • Classic River Crossing Problem Failure: However, when presented with a modified river crossing problem (where only the goat needed to be transported), the model reverted to solving the classic version, outlining steps to transport all items. This demonstrates a failure to adhere to the specific, modified constraints of the problem.

2. Failures in Complex Coding and Creative Prompts: The model struggled with more complex coding and creative tasks.

  • TV Channel Coder: A prompt to code a TV channel with channel switching (0-9), a channel name inspired by classic genres, and animations resulted in a blank screen with non-functional keys.
    • Comparison: The non-reasoning version of Quinn 3 Max produced a "pretty decent output" for this prompt.
  • 20 Bouncing Balls in a Heptagon: A demanding prompt involving 20 bouncing balls within a heptagon also failed. The model produced a simple animation, not adhering to the stricter requirements.
    • Comparison: The non-thinking version successfully generated code that adhered to the prompt, including a visually interesting heptagon and the ability to control rotation direction (clockwise/anticlockwise).

Technical Considerations and Observations

  • Model Size and Type: Quinn 3 Max is not an open-weight model, and its size is not publicly disclosed. It is positioned as one of Quinn's most powerful models.
  • One-Shot Testing Limitations: The reviewer acknowledges that one-shot testing for coding challenges might not be a definitive measure of agentic coding tool capabilities.
  • Generation Speed: The speed of code generation was similar between the thinking and non-thinking versions, with the difference primarily in the initial reasoning tokens.

Conclusion and Key Takeaways

Quinn 3 Max Thinking presents a complex picture. It excels in verifying information and demonstrating advanced reasoning, particularly with its interleaved tool-calling capabilities. However, it suffers from misguided attention in certain reasoning scenarios and has shown significant failures in more intricate coding and creative tasks, where the non-thinking version of Quinn 3 Max appears to perform better.

The reviewer suggests that for programming tasks, the non-reasoning version of Quinn 3 Max might be preferable based on their testing. They encourage users to share their own experiences with the model. Quinn 3 Max Thinking is available for free on the Quinn Chat platform.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video