MiniMax M2

MiniMax llm

A 230B-parameter (10B active) MoE model that deliberately returned to full attention over M1's linear attention, optimized for coding and agentic workflows.

overall ⌄

34.4

76 of 93 ranked

price · $/M tokens

0.3 / 1.2

in / out, as of 2026-07-25

to run it yourself ⌄

138 GB

minimum · 552 GB recommended

Benchmarks

What its maker published, and what anyone else measured. Every figure links the document it came from.

benchmarkpublishedmeasured
MMLU-Pro82
GPQA78
AIME 202578
SWE-bench69.4 61
HLE12.5
BrowseComp44

Cheaper, and at least as good

On a 1,000-in, 1,000-out call these cost less and rank no lower — a hard comparison to argue with, and a narrow one. Value covers what it misses.

Devstral Small 2 Mistral AI$0.00040overall 42.7
Ministral 3 14B Reasoning Mistral AI$0.00040overall 48.1
DeepSeek V4-Flash DeepSeek$0.00042overall 59.3
Qwen3 32B Alibaba$0.00080overall 35.1
DeepSeek V4-Pro DeepSeek$0.00130overall 68.3

Lineage

↓ finetune MiniMax M2.1
view in full graph →

Metadata

orgMiniMax
released2025-10-27
params230B-A10B
licensemodified-mit
hf idMiniMaxAI/MiniMax-M2