Minimax M2.1 : RIP Gemini 3.0 FLASH! This is the BEST SMALL & CHEAP OPEN MODEL YET!

AICodeKingAbout 4 min readDec 26, 2025Watch original
THE SUMMARYAI-generated

Miniaax M2 2.1 Model Review & Benchmarks

Key Concepts:

  • Miniaax M2 2.1: A new version of the Miniaax M2 language model, focused on agentic tasks and coding.
  • Agentic Benchmarks: Tests evaluating a model’s ability to perform complex tasks autonomously, often involving tool use and iterative problem-solving.
  • Non-Agentic Benchmarks: Standard tests assessing a model’s capabilities in areas like image generation, code creation (without autonomous execution), and basic reasoning.
  • Gemini Flash/3.0: Google’s language models, used as a performance and cost comparison point.
  • Open Weights: The model’s parameters are publicly available, allowing for local hosting and customization.
  • Claude Code/Kilo Code: Code editing environments used for testing agentic capabilities.
  • Expo/Godot/Tari/Nux: Frameworks and engines used in agentic benchmark tasks (app development, game creation).

I. Introduction & Model Overview

The video presents a review of Miniaax M2 version 2.1, a language model that builds upon the strengths of its predecessor (M2), particularly in agentic tasks. The reviewer received early access to the model and conducted a series of benchmarks to assess its performance. Miniaax M2 2.1 is expected to be generally available soon, priced at $120 per million input tokens and $30 per million output tokens – approximately half the cost of Gemini Flash. The reviewer is also currently testing another upcoming model (details withheld) which also shows promising results.

II. Non-Agentic Benchmark Results

The reviewer conducted several non-agentic benchmarks to evaluate the model’s core capabilities:

  • Floor Plan Generation: Performance was deemed “not great,” producing a functional but illogical floor plan.
  • SVG Image Generation (Panda with Burger): The result was “kind of fine,” with a minor suggestion for improved background separation.
  • Pokéball in 3D (3J’s): Achieved a perfect score (20/20) due to its precise dimensions and slick execution.
  • Chessboard with Autoplay: Failed to function correctly despite multiple attempts.
  • Minecraft Clone (Three.js, Kandinsky Style): Successfully generated a Minecraft-style clone adhering to the requested Kandinsky aesthetic, with accurate gravity simulation.
  • Majestic Butterfly in Garden: Successfully generated a visually appealing image with detailed wings and eyes.
  • CLI Tool in Rust: Generated a functional and well-structured CLI tool.
  • Pokéball in Blender Script: Considered a “major downer” and failed to produce a working script.
  • Math Questions: Failed to solve math questions, except for a riddle.

These benchmarks resulted in an overall ranking of 12th on the leaderboard, placing it above GLM4.6 and comparable to the more expensive 4.1 Opus model. In textual benchmarks, Miniaax M2 2.1 scored 53% compared to Gemini 3.0 Flash’s 47%, demonstrating superior performance at a lower cost.

III. Agentic Benchmark Results

The reviewer then focused on agentic benchmarks, assessing the model’s ability to autonomously create and execute code:

  • Go Tui Calculator: Successfully generated a functional and visually appealing calculator in a single attempt, utilizing Claude Code as the testing environment (with potential future compatibility with Kilo Code).
  • Movie Tracker App (Expo): Generated a functional homepage and search functionality, but inner pages lacked functionality. Still considered a “solid one-shot generation” given the model’s size and cost.
  • Godot Game: Successfully created a working game, highlighting increasing support for Godot within language models.
  • Tari App & Nux App: Both failed to generate functional applications, a common limitation even for more powerful models like Sonnet.
  • Open Code Repo: Generated a design but lacked functional code.

These agentic benchmarks resulted in an 8th position ranking, surpassing Gemini Flash with anti-gravity. The reviewer expressed a preference for Miniaax M2 2.1 over Gemini 3 Flash for most agentic tasks due to its lower cost, relative speed, and potential for enhancement through combination with models like Opus. The model’s small size also allows for potential local hosting with a suitable AI server.

IV. Key Arguments & Perspectives

The central argument is that Miniaax M2 2.1 represents a significant value proposition in the landscape of language models. It offers competitive performance, particularly in agentic tasks, at a substantially lower cost than alternatives like Gemini Flash. The reviewer emphasizes the importance of open weights, enabling local hosting and customization.

Notable Quote: “I really prefer this model for most agentic tasks over Gemini 3 Flash. It is much cheaper, relatively faster.”

V. Technical Details & Vocabulary

  • Tokens: Units of text used for processing by language models. Pricing is typically based on the number of input and output tokens.
  • Three.js: A JavaScript library for creating 3D graphics in a web browser.
  • Kandinsky Style: A style of art characterized by abstract forms and vibrant colors.
  • Expo: A framework for building native mobile apps with JavaScript.
  • Godot: A free and open-source game engine.
  • Rust: A systems programming language known for its safety and performance.
  • CLI: Command Line Interface.
  • SVG: Scalable Vector Graphics.

VI. Conclusion

Miniaax M2 2.1 is a promising language model, particularly for users focused on agentic tasks and AI coding. Its affordability, combined with strong performance and the benefit of open weights, makes it a compelling alternative to more expensive models. The reviewer anticipates the release of another even more impressive model in the near future. The model’s ability to perform well on complex tasks with limited resources positions it as a valuable tool for developers and researchers alike.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.