Minimax M2 (Fully Tested): I am switching to this. Better than Claude & GLM-4.6 on Long Running Task

By AICodeKing

Share:

Key Concepts

  • Minimax M2: A new, compact, high-efficiency large language model from Minimax, optimized for end-to-end coding and agentic workflows.
  • Activated Parameters: The subset of parameters in a neural network that are actively used during inference. Minimax M2 has 10 billion activated parameters out of 230 billion total.
  • Agentic Workflows: Tasks that involve an AI agent interacting with its environment, using tools, and performing multi-step actions to achieve a goal.
  • Tool Calling: The ability of an LLM to identify and utilize external tools (like APIs or command-line interfaces) to perform specific actions.
  • Context Window: The amount of text (measured in tokens) that a language model can consider at any given time.
  • Hugging Face: A platform that hosts and shares AI models, datasets, and code.
  • Open Router: A platform that provides API access to various LLMs.
  • Kilo: A framework or tool used for agentic tasks, which can be configured with LLM APIs.
  • Artificial Analysis Benchmarks: A set of benchmarks used to evaluate LLM performance, though the speaker expresses skepticism about their utility for saturated open benchmarks.
  • Custom Benchmarks: The speaker's own set of tests designed to evaluate specific capabilities of LLMs, including creative generation, coding, and agentic task execution.

Minimax M2: A Detailed Analysis

This video provides an in-depth review of the newly released Minimax M2 model, comparing its performance against other leading LLMs, particularly in coding and agentic tasks.

Model Overview and Availability

  • Model Name: Minimax M2
  • Previous Version: Minimax M1
  • Availability: Weights are available on Hugging Face, suggesting potential open-sourcing.
  • Current Access: Free on Open Router and Minimax's API platform. This allows integration with tools like Kilo for agentic workflows.
  • API Details: Minimax's own API might offer better rate limits than Open Router. The model is designed for extensive use in various coding scenarios.

Technical Specifications and Performance Metrics

  • Model Size: Compact and efficient, with 230 billion total parameters and 10 billion activated parameters. This is significantly smaller than models like GLM 4.5 Air (110 billion parameters smaller).
  • Capabilities: Optimized for end-to-end coding and agentic workflows. Delivers "near-frontier intelligence" in general reasoning, tool use, and multi-step task execution.
  • Key Features: Low latency and deployment efficiency.
  • Artificial Analysis Benchmarks:
    • Scores just below Claude 4.5 Sonnet.
    • Speed is considered "pretty fine."
    • Pricing: $0.5 per million tokens (input) and $2.2 per million tokens (output).
    • Context Window: 205,000 tokens. This is a notable reduction from the previous model's 1 million tokens, a decision the speaker questions.
    • Coding Index: Scores two points below Sonnet. The speaker expresses doubt about the validity of these benchmarks, citing Grok 4 Fast's high score despite being a poor coding model.
  • Speaker's Custom Benchmarks:
    • Floor Plan Question: The model generates a floor plan, but it lacks logical coherence. Scored accordingly.
    • Panda Holding a Burger: Performed "pretty good," ranking among the best open models, though not as strong as Gemini 3's checkpoints.
    • Pokeball in 3JS: Not good; resembled a Premier Ball rather than a Pokeball.
    • Chessboard: Laid out correctly but did not function. The speaker suspects training on GPT-5 outputs due to the UI similarity.
    • Minecraft Game: Did not work.
    • Butterfly Flying in the Garden: "Kind of fine," resembling a bug but functional.
    • CLI Tool in Rust & Blender Script: "Fine, but not great."
    • Mathematics Question: Passed one question.
    • Riddle Question: Passed.
    • Overall Custom Benchmark Ranking: Placed 12th on the speaker's leaderboard, below Claude 4.5 Sonnet, GLM, and Deepseek Terminus. The speaker highlights that Minimax M2, GLM, and LongCat are the only top-15 models of this size, making Minimax M2's performance an "awesome win" given its small parameter count.

Agentic Task Performance

The speaker emphasizes that Minimax M2 truly shines in agentic tasks, calling it a "true agentic model."

  • Integration with Kilo: Works "extremely well" with Kilo for agentic tests, configurable via Minimax M2 API or Open Router. The speaker uses the Open Router API.
  • Edit Failures: The first open model observed that does not produce edit failures in agentic tasks.
  • Movie Tracker App: "Really good." Features sliding panels and inner page navigation. A minor drawback is the title bar not being removed.
  • Code Quality: "Insane." Avoids common issues like hardcoding API keys (unlike Sonnet). Organizes code into different files for better management.
  • GUI Calculator App: "Pretty great." Works well and effectively utilizes all tools within Kilo Code, including search and replace and terminal command execution.
  • Godo Game: "Just not good." The model struggles with the Godo language. The speaker acknowledges this is acceptable given the model's size and cost.
  • Open Code Repo Question (Go): The model navigated files correctly, which is a challenge, but the overall task was not completed successfully. The speaker notes that even Sonnet struggles with this.
  • Spelt Question: "Kind of fine," reaching a "somewhat usable" point.
  • Long-Running Tasks: Performs "quite good" on long tasks, outperforming GLM 4.6, which starts to falter. Minimax M2 can run for "hours," similar to GPT-5.
  • Coding (General): Not a strong suit.
  • Rust: Not a strong spot.
  • Agentic Task Leaderboard Ranking: Fifth position.
  • Comparison to GLM 4.6: Slightly below GLM 4.6 in agentic tasks but considered superior for general use cases due to its long-running task capabilities.

Key Arguments and Perspectives

  • Skepticism towards Benchmarks: The speaker expresses a lack of confidence in standard benchmarks for evaluating LLMs, especially those that are saturated and potentially overfitted by models.
  • Value Proposition of Minimax M2: The model offers exceptional performance for its size and cost, particularly in agentic workflows and long-running tasks.
  • Potential for Replacement: The speaker suggests Minimax M2 might be a viable alternative to GLM, despite GLM's coding advantages, due to its efficiency and cost-effectiveness.
  • "Awesome Win" for Minimax: The model's performance, especially in agentic tasks, is considered a significant achievement for Minimax, given its compact size compared to competitors.

Notable Quotes

  • "This is a pretty small model. It's only 230 billion parameters with about 10 billion activated parameters."
  • "I don't really find these benchmarks useful at all apart from the speed or provider variance benchmarks because they mainly use open benchmarks that are very saturated at this point."
  • "This is an insanely small model when compared to GLM or DeepSeek. So, this is an awesome win for them."
  • "I mean this is a true agentic model."
  • "This is the first open model I've seen that doesn't give me any edit failures. It's just really good at agentic tasks."
  • "The code quality of this model is insane."
  • "It can keep going for hours, similar to GPT5, and that is an awesome thing."

Conclusion and Future Outlook

Minimax M2 is presented as a highly impressive and efficient large language model, particularly excelling in agentic tasks and long-running operations. Its compact size and low cost make it a compelling option for developers and AI enthusiasts. While it has limitations in certain coding scenarios and specific creative generation tasks, its strengths in agentic workflows and overall efficiency position it as a strong contender in the LLM landscape. The speaker plans to conduct further testing and potentially release another video detailing more nuances of the model.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video