← Log
5 min read

Most models are beaten outright

A ranking is a line. It can tell you that one model is better than another and nothing else — not what the difference costs, not whether it was worth paying for. Put cost on a second axis and the picture changes completely, because on a plane a model can be beaten on both counts at once.

Run that over every model in the catalog with a published price and a standing in the table, and the result is worse than expected.

43 of 52

priced, ranked models that something else beats on price and standing at the same time

9

left on the frontier — every defensible choice is one of these

150×

between the cheapest and the dearest model that is still worth considering

Eighty-three percent of the field is dominated. For each of those models there exists another that costs no more, scores no less, and beats it on at least one of the two. Not "better value in my opinion" — beaten on both axes, arithmetically.

What domination means, precisely

  1. pick a model
  2. find anything cheaper
  3. check its standing
  4. if it is also no lower, the first model is beaten

The rule is deliberately strict. A tie on both axes is not domination — two models at the same price with the same standing are both defensible picks, and the frontier keeps both. Only a strict win on one axis with no loss on the other counts.

That strictness matters, because it means being dominated is not a matter of taste. There is no weighting to argue about, no composite score with a thumb on the scale. Either something else is cheaper and at least as good, or it is not.

The plane

better ↑ dearer → Claude Mythos 5 Claude Opus 5 DeepSeek V4-Flash DeepSeek V4-Pro Gemini 3.1 Pro
cost per call (log scale) overall
Every priced, ranked model. Cost of a 1,000-in, 1,000-out call runs left to right on a log scale — the field spans three orders of magnitude — and standing runs bottom to top. Anything up and to the left of a dot beats that dot outright.

The shape is the argument. A ranking would have given you a single column; this is a plane with most of its occupants strictly inside the boundary, not on it.

Walking the frontier

Nine models survive. Read them in cost order and each row is a decision: this is what the next step costs, and this is what it buys.

modelper calloverallwhat the step buys
DeepSeek V4-Flash$0.000459.8
DeepSeek V4-Pro$0.001369.1+9.2 for +$0.0009
Gemini 3.1 Pro$0.014072.6+3.5 for +$0.0127
GPT-5.2$0.015873.7+1.1 for +$0.0017
Claude Sonnet 5$0.018074.0+0.3 for +$0.0022
Claude Opus 5$0.030081.7+7.7 for +$0.0120
Claude Mythos 5$0.060084.2+2.6 for +$0.0300
The top of the frontier, from DeepSeek V4-Flash upward. Two cheaper entries — Devstral Small 2 and Ministral 3 14B Reasoning — sit below it at the same $0.0004 with lower standing.

The highlighted row is where the money is. Nine hundred microdollars buys nine points of standing — the best trade anywhere on the curve.

And then the curve breaks:

V4-Flash → V4-Pro: +9.2 pts 3%
V4-Pro → Gemini 3.1 Pro: +3.5 pts 42%
Gemini → GPT-5.2: +1.1 pts 6%
Opus 5 → Mythos 5: +2.6 pts 100%
What each step up the frontier costs, as a share of the largest. The third step is where the price stops tracking the capability.

Going from DeepSeek V4-Pro to Gemini 3.1 Pro costs eleven times more per call and buys 3.5 points. That is the cliff — the point where you stop paying for capability and start paying for something else. It may well be worth it: the something else could be reliability, latency, a support contract, or a benchmark that matters to you and does not appear in the overall column. But it should be a decision, and on a ranking it is invisible.

The expensive dominated models

The dominated set is not made up of obscure models nobody uses. The three most expensive calls in the entire catalog are all dominated:

modelper calloverallbeaten by
Claude Opus 4$0.090041.2Devstral Small 2 — 225× cheaper, higher standing
o1$0.075036.0everything on the frontier
Claude Fable 5$0.060073.2Claude Mythos 5 — same price, 11 points higher
Two are previous-generation models still on their vendors' price lists; the third is beaten by a sibling at an identical rate.

Claude Opus 4 is the extreme: the most expensive call in the table, at a standing below a model that costs a two-hundred-and-twenty-fifth of it. It is a deprecated model still published at its old price, which is exactly the pattern to watch for — price lists are sticky and capability is not. Nobody at a vendor lowers the price of last year's flagship when this year's is better and cheaper; they just stop mentioning it.

The Fable 5 case is different and sharper. Same maker, same price as Mythos 5, eleven points of standing behind it. Nothing about that is a mistake — the two serve different purposes — but if the only two axes you care about are cost and general standing, the choice is made for you.

What this does not tell you

The plane has two axes and your decision has more.

Dominated does not mean bad. It means that on these two measures something else wins on both. A dominated model can still be the right pick for a reason the plane cannot see: a context window, a licence, where it is hosted, a latency profile, a data-residency requirement, an existing contract.

Overall is general and your job is specific. It averages a model's standing across the benchmarks it publishes, discounted for how few of them there are. If you are shipping a coding agent, the SWE-bench column is a better guide than this one, and the frontier would look different if you drew it against a single benchmark rather than the average.

One reference call is a simplification. Everything here is priced on 1,000 tokens in and 1,000 out, which makes the comparison fair and your workload not. A retrieval-heavy pipeline sending fifty thousand input tokens for a two-hundred-token answer has a completely different shape, and the models with cheap input move up. The token cost page prices your actual prompt.

Prices go stale, quietly. Every figure is a vendor list rate on the day it was recorded, and the Opus 4 case shows what happens when nobody revisits one.

The habit worth taking

Before paying for a model, look left. Everything to the left of it on that plane costs less; if any of them sits at or above its height, you are paying for something that is not capability — and it might be the right call, but you should be able to say what it is.

The whole plane is at Value, and every model page now says outright whether anything beats it.