M

MiniMax: minimax-m2-her

minimax-m2-her
204.8K context2.0K outReasoningToolsCache
Released Jan 23, 2026Knowledge cutoff Jun 2025Updated Sep 24, 2026

MiniMax-M2 is a high-efficiency Mixture-of-Experts (MoE) model architected specifically for coding and agentic workflows. While boasting 230B total parameters, it only activates 10B per token, ensuring low latency and high throughput. Its standout feature is "Interleaved Thinking," where the model uses internal reasoning tags (<think>) to plan and self-correct during complex tasks. This makes it exceptionally robust for multi-step tool use, terminal-based coding, and long-horizon planning. It ranks as a top-tier open-weight model, rivaling proprietary giants in agentic benchmarks like VIBE and SWE-bench.

Mode chatTokenizer Other

Pricing

Input price
$0.60/ 1M tokens
Output price
$2.40/ 1M tokens
Context window 204.8K tokensCompatible endpoints openaiVendor MiniMax

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltySome providersNot sent by default
max_tokensAll providers-
temperatureAll providers1
top_pAll providers0.95

Frequently asked questions

Similar models