M

MiniMax: minimax-m3

minimax-m3
1M context512K outReasoningToolsVisionVideoCacheStructured
Released May 31, 2026Updated Jun 18, 2026

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks. Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

Mode chatTokenizer OtherQuantization fp8MiniMaxAI/Minimax-M3

Pricing

Input price
$0.37/ 1M tokens
Output price
$1.49/ 1M tokens
Compatible endpoints openaiVendor MiniMax

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providersNot sent by default
include_reasoningAll providers-
logit_biasSome providers-
logprobsSome providers-
max_tokensAll providers-
min_pSome providers-
presence_penaltyAll providersNot sent by default
reasoningAll providers-
repetition_penaltyAll providersNot sent by default
response_formatAll providers-
seedSome providers-
stopAll providers-
structured_outputsSome providers-
temperatureAll providers1
tool_choiceAll providers-
toolsAll providers-
top_kAll providersNot sent by default
top_logprobsSome providers-
top_pAll providers0.95

Frequently asked questions

Similar models