Value
A ranking is a line — it can only say which model is better. Cost against standing is a plane, and on a plane most models are beaten outright: something else is cheaper and at least as good.
44 of 52
priced, ranked models that something else beats on both price and standing
8
left on the frontier — every defensible choice is one of these
150×
between the cheapest and dearest model on the frontier
The plane
One dot per model that has both a published price and a standing in the table. Cost runs left to right on a log scale — the field spans three orders of magnitude — and standing runs bottom to top. Anything up and to the left of a dot beats it outright. The lit dots are the frontier.
The frontier
Read top to bottom: each row costs more than the one above it and buys something for the difference. Nothing between two rows is worth considering, because anything there is beaten by one of them.
| model | per call | overall | the step ⌄ |
|---|---|---|---|
| Devstral Small 2 Mistral AI | $0.00040 | 42.7 | — |
| Ministral 3 14B Reasoning Mistral AI | $0.00040 | 48.1 | +5.5 for +0.00000 |
| DeepSeek V4-Flash DeepSeek | $0.00042 | 59.3 | +11.2 for +0.00002 |
| DeepSeek V4-Pro DeepSeek | $0.00130 | 68.3 | +9.0 for +0.00088 |
| Gemini 3.1 Pro Google | $0.0140 | 71.6 | +3.4 for +0.0127 |
| GPT-5.2 OpenAI | $0.0158 | 73.7 | +2.0 for +0.00175 |
| Claude Opus 5 Anthropic | $0.0300 | 81.7 | +8.1 for +0.0143 |
| Claude Mythos 5 Anthropic | $0.0600 | 83.7 | +2.0 for +0.0300 |
how to read this
- Dominated does not mean bad. It means that on these two axes something else wins on both. A dominated model can still be the right pick for a reason this plane cannot see — a context window, a licence, a latency profile, where it is hosted.
- Overall is a general figure, and your job is specific. It averages a model's standing across the benchmarks it publishes. A coding workload should be reading the SWE-bench column, not this one; the frontier is a starting point for the question rather than its answer.
- One reference call, so the comparison is fair. Every model is costed on the same 1,000 tokens in and 1,000 out. For your own prompt, use Token cost.
- Only models with a first-party price appear. Open weights with no API of their own have no price to plot — what they cost is your own hardware.