MiniMax M2.5 (Fully Tested): I've been testing it for the last 4 days and it is AMAZING!!!

AICodeKingAbout 4 min readFeb 14, 2026Watch original
THE SUMMARYAI-generated

Miniax M2.5 Model Evaluation

Key Concepts:

  • Miniax M2.5: A new frontier-level language model from Miniax, positioned as a cost-effective alternative to leading models like Opus and Gemini.
  • Frontier Model: A high-performance language model capable of complex tasks, typically requiring significant computational resources.
  • Agentic Applications: Utilizing language models to create autonomous agents capable of performing tasks, such as coding, data analysis, or application development.
  • Tokens: The basic units of text processed by language models; cost is often calculated per token.
  • TPS (Tokens Per Second): A measure of the model’s processing speed.
  • Catching: A feature allowing the model to correct its own errors during generation.
  • M2.5-Lightning: A faster version of M2.5, prioritizing speed over cost.

1. Model Overview & Cost Analysis

The Miniax M2.5 is presented as a successor to the M2 and M2.1 models, maintaining the same parameter size (230 billion) while significantly improving performance and reducing cost. A core selling point is its affordability, aiming to deliver “intelligence too cheap to meter.” The model is available in two versions: M2.5 and M2.5-Lightning.

  • M2.5: Operates at 50 tokens per second (TPS) and costs $0.30 per million input tokens and $2.40 per million output tokens. Running it continuously for an hour at 50 TPS costs $0.30.
  • M2.5-Lightning: Operates at 100 TPS, doubling the speed, and costs $0.30 per million input tokens and $2.40 per million output tokens. Continuous operation for an hour at 100 TPS costs $1.
  • Cost Comparison: Miniax claims M2.5 is 1/10th to 1/120th the cost of Opus, Gemini 3 Pro, and GPT-5. Four instances of M2.5 can run continuously for a year for $10,000.

2. Performance & Benchmarking

The presenter emphasizes that M2.5 is designed for agentic applications, not general-purpose benchmarks. Despite this, it demonstrates performance comparable to, and sometimes exceeding, Opus 4.6, a significantly more expensive model. Specifically, it’s noted to be 30 times cheaper than Opus while offering comparable performance. The model’s ability to “think through” errors and self-correct (“inner thinking”) is highlighted as a key strength.

3. Agentic Application Testing & Results

The presenter tested M2.5 across several agentic applications, utilizing the Kilo CLI (a fork of OpenCode with enhanced features) and Kilo Gateway. The results were consistently positive:

  • Movie Tracker App (Expo): Completed in approximately 4 minutes, outputting 52,000 tokens. The model successfully implemented features like movie lists, reviews, and a calendar. While the UI wasn’t a primary focus, the functionality was described as “otherworldly” given the model’s size and cost.
  • Calculator in Terminal (Go & Bubble Tea): Successfully built a functional calculator, including checking for Go installation and installing necessary libraries. The layout was described as compact and well-organized.
  • Tori Desktop App (Image Cropper): Successfully created an image cropper tool, a task that even Opus struggles with. The tool included features like aspect ratio control and a toolbar.
  • Stack Overflow Clone (Nux App, Database & O): Created a functional Stack Overflow clone with threads and a database.
  • Spelt Conbound App: Successfully built a task management application with boards, lists, and tasks.

4. Comparative Analysis with Gemini 3 Flash

A direct comparison was made with Gemini 3 Flash, revealing that Miniax M2.5 costs $1.20 for output while Gemini 3 Flash costs $3 for output, with Miniax being superior in coding performance.

5. Release Timing & Context

The timing of the M2.5 release is linked to the upcoming Chinese New Year, with AI labs aiming to deploy models before the holiday. The presenter describes this as a “wholesome” trend.

6. Leaderboard Ranking & Overall Assessment

M2.5 currently ranks fourth on the presenter’s agentic application leaderboard. The overall assessment is highly positive, emphasizing the model’s speed, cost-effectiveness, and impressive performance in agentic tasks.

Notable Quote:

“It’s so good. That is actually amazing.” – The presenter, expressing enthusiasm for the model’s performance.

7. Technical Terms Explained:

  • Parameter Size: The number of variables a model learns during training; generally, larger parameter sizes indicate greater model capacity. (M2.5 has 230 billion parameters)
  • OpenCode/Kilo CLI: Tools used for building and interacting with agentic applications.
  • Expo/Go/Bubble Tea/Nux: Specific technologies and frameworks used in the tested applications.

Conclusion:

Miniax M2.5 represents a significant advancement in accessible AI, offering frontier-level performance at a dramatically lower cost than competing models. Its strength lies in agentic applications, where it demonstrates impressive capabilities in coding, application development, and task automation. The model’s speed, affordability, and self-correcting abilities position it as a compelling option for developers and researchers seeking to build innovative AI-powered agents. The release of both standard and “Lightning” versions provides flexibility based on specific performance and cost requirements.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.