Benchmarks

What each one measures, and how the catalog ranks on it. Scores are self- or third-party reported — every number links its source on the model page.

Multi-subject knowledge and reasoning across 57 academic tasks (5-shot accuracy).

1Llama 3.1 405B Instruct88.6
2Llama 3.1 405B85.2
3Llama 3.1 8B66.7

Graduate-level, Google-proof science questions — hard reasoning signal.

1DeepSeek R171.5
2Llama 3.1 405B Instruct51.1
3DeepSeek R1 Distill Llama 8B49

SWE-bench

about ↗

Resolving real GitHub issues end-to-end in real repositories.

No catalog models scored yet.