Every model. Every benchmark. One table.
Curated catalog with lineage and sourced scores.
112 models · 7 benchmarks
run
columns
how to read this
- Every figure is published by the model's maker. We run no benchmarks of our own. Each score links to the model card, paper or announcement it was taken from — follow it before you rely on it.
- An empty cell means unpublished, not zero. Makers report the benchmarks that suit them, so a sparse row is a claim not made rather than a claim failed.
- Overall ranks each model within every column it enters, then averages those standings and pulls thin evidence toward the middle: one reported benchmark is a weaker claim than five, and the number says so.
- A figure in this colour was measured by someone else. Independent platforms run the evaluation themselves, the same way for every model, and publish what they got. Where one sits beside a maker's own number, the gap between them is the interesting part; where it sits alone, it is the only measurement anyone has published. Neither is automatically the truer figure — but the ranking is built only from what makers publish, so a measured figure never moves a model up the table.
- A score is also a claim about effort. The same model answers differently at max reasoning and at its default setting, and makers headline whichever suits them. Where the source states the setting it is marked on the figure; where it does not, the tooltip says so rather than pretending the run was standard.
- Compare down a column, not across the table. Makers evaluate on their own harnesses, and two labs reporting the same benchmark have not necessarily run the same test.