Every model. Every benchmark. One table.

Curated catalog with lineage and sourced scores.

112 models · 7 benchmarks

run
columns
Claude Mythos 5 84 $10 $5094.1 95.5 thinking 80.3 thinking 59 auto 88 max
Claude Opus 5 82 $5 $2593.7 96 79.2 56.3 52.6 90.8
Claude Opus 4.8 78 $5 $2593.6 88.6 69.2 49.8 84.3
GPT-5.6 77 $5 $3094.6 94.1 64.6 90.4
Claude Opus 4.7 75 $5 $2594.2 87.6 thinking 64.3 thinking 46.9 max 79.3 max
GPT-5.2 74 $1.75 $1492.4 100 80 71.8 69
Claude Fable 5 73 $10 $5095 thinking 80 thinking 53.3
Claude Sonnet 5 73 $3 $1585.2 thinking 63.2 thinking 43.2 auto 84.7 max
Gemini 3.1 Pro 72 $2 $1294.3 94.1 80.6 54.2 44.4 85.9
Claude Opus 4.6 69 $5 $2591.31 99.79 80.84 75.6 53.4 40 83.7
DeepSeek V4-Pro 68 $0.435 $0.8787.5 90.1 80.6 55.4 37.7 83.4
GLM-5.2 68 $1.4 $4.491.2 62.1 40.5
Hy367 self-host self-host90.4 78 57.9 53.2
GPT-5.5 67 $5 $3093.6 59.4
Kimi K2.6 65 $0.95 $490.5 80.2 58.6 36.4 83.2
MiniMax M3 65 $0.3 $1.280.5 59 83.5
Claude Sonnet 4.6 64 $3 $1589.9 95.6 79.6 58.1 34.6 74
Qwen3.5 397B A17B 64 $0.6 $3.687.8 88.4 76.4 69
GPT-5.4 64 $2.5 $1592.8
GPT-5.1 62 $1.25 $1088.1 94 76.3 66
Claude Opus 4.5 60 $5 $2587 92.77 80.9 76.8 52
Qwen3.6-27B59 self-host self-host86.2 thinking 87.8 thinking 77.2 thinking 53.5 thinking
DeepSeek V4-Flash 59 $0.14 $0.2886.2 88.1 79 52.6 34.8 73.2
Kimi K2.5 59 $0.6 $387.1 thinking 87.6 thinking 96.1 thinking 76.8 70.8 50.7 31.5 60.6
GPT-5 59 $1.25 $1094.6 74.9 65
Gemma 4 31B Instruct58 self-host self-host85.2 84.3
GPT-5.1-Codex-Max 57 $1.25 $1077.9
Gemini 3.6 Flash 56 $1.5 $7.558.7
Mistral Medium 3.556 $1.5 $7.577.6
Claude Sonnet 4.5 54 $3 $1577.2 71.4 70.6
Gemini 3.5 Flash 54 $1.5 $953.9 40.2
Gemini 3 Flash 54 $0.5 $390.4 78 75.8 48.4 33.7
GLM-5.1 53 $1.4 $4.486.2 58.4 31 68
GLM-4.7 53 $0.6 $2.284.3 85.7 95.7 73.8 24.8 52
GLM-5 53 $1 $3.286 77.8 72.8 30.5 62
Gemma 4 26B A4B Instruct52 self-host self-host82.6 82.3
MiniMax M2.1 51 $0.3 $1.288 83 83 74 22.2 47.4
Gemini 3 Pro 50 91.9 76.2 69.6 43.3 37.5 59.2
DeepSeek V3.2-Speciale 49 self-host self-host30.6
Gemini 3.5 Flash-Lite 49 $0.3 $2.554.2
Ministral 3 14B Reasoning 48 $0.2 $0.285
Nemotron 3 Ultra 550B A55B 47 self-host self-host86.8 70.7 26.7 44.4
Qwen3.6-35B-A3B47 self-host self-host85.2 thinking 86 thinking 73.4 thinking 49.5 thinking 21.4 thinking
DeepSeek R1-0528 47 self-host self-host85 81 87.5 57.6
GLM-4.546 $0.6 $2.284.6 79.1 64.2 54.2
Devstral 246 $0.4 $272.2
DeepSeek V3.2 46 self-host self-host73.1 70 25.1
Grok 345 79.9 75.4
DeepSeek R1 44 self-host self-host71.5
Gemini 2.5 Pro 44 $1.25 $1086.4 88 59.6 53.6 21.6
Ornith-1.0-35B44 self-host self-host75.6 50.4
Gemma 4 12B Instruct43 self-host self-host77.2 78.8
Claude Haiku 4.5 43 $1 $573 80.7 73.3 66.6
Devstral Small 243 $0.1 $0.368
DeepSeek V3.2-Exp 42 self-host self-host85 79.9 89.3 67.8 19.8 40.1
Llama 4 Maverick Instruct42 self-host self-host80.5 69.8 21
Qwen3 235B A22B 42 $0.7 $2.871.1 81.5
Claude Opus 4 41 $15 $7579.6 thinking 75.5 thinking 72.5 none 67.6
Mistral Large 341 $0.5 $1.567.17
DeepSeek R1-Distill-Qwen-32B 40 self-host self-host62.1
Claude 3.5 Sonnet39 59.4
Nemotron 3 Super 120B A12B 38 self-host self-host83.73 90.21 60.47 18.26 31.28
Gemini 2.0 Flash38 77.6 60.1 13.5
Claude Sonnet 4 37 $3 $1575.4 thinking 70.5 thinking 72.7 none 64.9
Llama 3.1 405B Instruct37 self-host self-host51.1
Olmo 3-Think 32B 37 self-host self-host72.5
Gemini 3.1 Flash-Lite 36 $0.25 $1.586.9 38.3 16
Qwen3-Coder-Next36 self-host self-host70.6 44.3
o1 36 $15 $6078.3 40.9
Claude 3.7 Sonnet 36 68 63.7 52.8
DeepSeek R1 Distill Llama 8B 36 self-host self-host49
MiniMax M1 36 self-host self-host81.1 70 76.9 56
Qwen3 32B 35 $0.16 $0.6468.4 72.9
Devstral Small 1.135 self-host self-host53.6
DeepSeek V3.1 35 self-host self-host80.1 88.4 66 15.9 30
MiniMax M2 34 $0.3 $1.282 78 78 69.4 61 12.5 44
Ornith-1.0-9B34 self-host self-host69.4 42.9
GLM-4.7-Flash 34 self-host self-host91.6 59.2 14.4 42.8
GLM-4.634 $0.6 $2.268 55.4 17.2 45.1
Kimi K234 self-host self-host81.1 75.1 49.5 65.8
Llama 4 Scout Instruct34 self-host self-host74.3 57.2 9.1
Magistral Medium 34 $2 $570.8 64.9
GPT-4.132 $2 $866.3 54.6 39.6
Nemotron 3 Nano 30B A3B 32 self-host self-host78.3 89.1 38.8 10.6
Qwen2.5 72B Instruct32 self-host self-host71.1 49
Llama 3.3 70B Instruct32 self-host self-host68.9 50.5
Gemini 2.5 Flash 31 $0.3 $2.580.8 75.6 54 28.7 11
Gemma 3 27B Instruct29 self-host self-host67.5 42.4
Llama 3.1 70B Instruct29 self-host self-host66.4 48
DeepSeek V329 self-host self-host75.9 59.1 42
Granite 4.1 30B28 self-host self-host64.09 45.76
Qwen2.5-Coder 32B Instruct26 self-host self-host62.3 41.8 9
Llama 3.1 8B Instruct25 self-host self-host48.3 31.8
Gemini 1.5 Pro 65.7 37.1 3.9
Gemini 3 Deep Think
Gemma 2 27Bself-host self-host
GPT-4o$2.5 $1074 52.6 21.6 2.8
GPT-5.2-Codex $1.75 $1489.9 33.5
GPT-5.3-Codex $1.75 $1491.5 39.9
Grok 4 86.6 87.7 92.7 23.9
Grok 4.1
Grok 4.1 Fast 85.4 74.3 85.3 63.7 89.3 34.3 17.6 5
Grok 4.5 $2 $693.1 40.3
Hermes 3 Llama 3.1 8Bself-host self-host
Kimi K2.7-Code $0.95 $489.6 32.8
Llama 3.1 405Bself-host self-host73.2 51.5 3 4.2
Llama 3.1 8Bself-host self-host47.6 25.9 4.3 5.1
Llama 3.2 3B Instructself-host self-host34.7 22.4 3.3 5.2
Mistral Large 2self-host self-host68.3 47.2 0 3.2
Mixtral 8x22B$2 $653.7 33.2 4.1
o3 $2 $885.3 82.7 88.3 58.4 20
o4-mini $1.1 $4.483.2 78.4 90.7 45 17.5

how to read this

  • Every figure is published by the model's maker. We run no benchmarks of our own. Each score links to the model card, paper or announcement it was taken from — follow it before you rely on it.
  • An empty cell means unpublished, not zero. Makers report the benchmarks that suit them, so a sparse row is a claim not made rather than a claim failed.
  • Overall ranks each model within every column it enters, then averages those standings and pulls thin evidence toward the middle: one reported benchmark is a weaker claim than five, and the number says so.
  • A figure in this colour was measured by someone else. Independent platforms run the evaluation themselves, the same way for every model, and publish what they got. Where one sits beside a maker's own number, the gap between them is the interesting part; where it sits alone, it is the only measurement anyone has published. Neither is automatically the truer figure — but the ranking is built only from what makers publish, so a measured figure never moves a model up the table.
  • A score is also a claim about effort. The same model answers differently at max reasoning and at its default setting, and makers headline whichever suits them. Where the source states the setting it is marked on the figure; where it does not, the tooltip says so rather than pretending the run was standard.
  • Compare down a column, not across the table. Makers evaluate on their own harnesses, and two labs reporting the same benchmark have not necessarily run the same test.