Benchmarks
What each one measures, and how the catalog ranks on it. Scores are self- or third-party reported — every number links its source on the model page.
MMLU
about ↗Multi-subject knowledge and reasoning across 57 academic tasks (5-shot accuracy).
| 1 | Llama 3.1 405B Instruct | 88.6 |
| 2 | Llama 3.1 405B | 85.2 |
| 3 | Llama 3.1 8B | 66.7 |
GPQA
about ↗Graduate-level, Google-proof science questions — hard reasoning signal.
| 1 | DeepSeek R1 | 71.5 |
| 2 | Llama 3.1 405B Instruct | 51.1 |
| 3 | DeepSeek R1 Distill Llama 8B | 49 |
SWE-bench
about ↗Resolving real GitHub issues end-to-end in real repositories.
No catalog models scored yet.