If another AI model is at least as good and costs no more, the one you are paying for is dominated. More precisely: a model is dominated when an alternative scores at least as high on the same quality test, costs no more for the same kind of work, and is better on at least one of the two. On those two measures, keeping a dominated model buys you nothing.
The catch is a third question the two numbers cannot answer: can the alternative still do your job? On Undominated.ai, the independent AI inference price index I built, that question changes a lot. Its methodology page reports that without a capability check, 61% of raw “better and cheaper” verdicts pointed at a model that could not do the current model’s job (59 of 97, measured on the 24 August 2026 catalogue).
Where does the word “dominated” come from?
From optimisation. One option dominates another when it is no worse on every criterion and better on at least one (definition). The options that nothing dominates form the frontier. For each of them, nothing cheaper scores as high.
For AI models, the two criteria are quality and price. On the board dated 23 September 2026, 123 of the 136 models with both an independent score and a standard price had an alternative that scored at least as high for no more money. Only 13 had none: that is the frontier. The 123 cost a median 6.7 times as much as the cheapest such alternative.
Those numbers move with every catalogue update. The live board is the source, not this paragraph.
What are the three checks behind the label?
A dominated verdict is only as good as three comparisons. Each one can fail in a predictable way.
| Check | The plain question | What goes wrong if you skip it |
|---|---|---|
| Same quality test | Were both models scored on the same test? | A higher number from a different benchmark proves nothing about the one you chose. |
| Same work, same price | Is the price for the same mix of input and output? | A model that is cheap for short chats can be dear for long documents. |
| Same job | Can the alternative take in and give back everything yours does? | ”Better and cheaper” points at a model that lacks the one capability you rely on. |
Quality. Undominated uses LMArena Elo: a score built from crowdsourced votes in which people compare the answers of two models side by side (Chatbot Arena paper). Scores carry noise. On this board, two scores need to differ by about 10.55 points before the difference is statistically clear (a 95% interval), so models inside that margin share a rank.
Price. AI providers bill per token, a small chunk of text, often a short word or part of a longer one. Input (what you send) and output (what comes back) have separate rates, so the price of a model depends on the work. The board prices every model for a workload you choose: Balanced (three tokens in for every token out), Summarise, Chat, Code gen, or Agentic. It also flags the traps: 69 models change their rate once a prompt passes a length threshold, and some bill reasoning tokens (the model’s hidden working) on a separate meter.
Capability. Call it the capability envelope: what a model can take in and give back. That means context length (how much text it reads at once), maximum output, input types such as images or audio, tool use, and reasoning. The board only calls a replacement clean when its envelope covers the original. Of the 123 verdicts on 23 September 2026, 109 were clean and 14 (11%) named what you would give up.
Two real verdicts, read carefully
Both come from the board’s Check pages, Balanced workload, catalogue dated 23 September 2026. They are snapshots, not advice.
Dominated at the same price. Claude Opus 4.6 is dominated by Claude Opus 5. Both cost $10.00 per million tokens on this workload. Opus 5 scores 1.9 points higher and covers everything Opus 4.6 can take in and give back. The saving: 0%.
Now read the gap. 1.9 points is well inside the 10.55-point margin, and the board itself shows both models sharing rank 1. So the verdict means “there is no reason to pick the older model at the same price”. It does not mean “the newer model is measurably better”. Switching still costs work: The Model Portfolio treats prompts as assets tied to one model, because a prompt tuned for one model can perform worse on another.
Better and cheaper, but not a drop-in. Gemini 3.8 Flash scores 5.0 points higher than Muse Spark 1.3 and costs 25% less ($1.50 against $2.00 per million tokens). The board still refuses to call it a replacement, because you would give up maximum output: Muse Spark 1.3 can return up to 943,718 tokens in one answer. On the leaderboard, the Gemini row also notes that reasoning tokens are billed on top, so the 25% gap is before that charge.
If your job never needs very long answers, that trade may be worth a test. If it does, the “better and cheaper” model is the wrong model.
Quiz: dominated or not?
A hypothetical current model, A, scores 1,400, costs $4 per million tokens on your workload, reads 200,000 tokens of context, and accepts text and images. Cover the last column and decide for each alternative.
| Alternative | Score | Price | Envelope | Verdict |
|---|---|---|---|---|
| B | 1,420 | $3 | 200K, text only | Not a replacement. Better and cheaper, but it loses image input. A trade. |
| C | 1,400 | $3 | 200K, text and images | Dominates A. Same score, cheaper, same envelope. |
| D | 1,460 | $6 | 1M, text and images | Does not dominate A. Better, but dearer. A choice about money. |
| E | Unrated | $1 | 200K, text and images | Unknown. No independent score, so no quality verdict either way. |
| F | 1,404 | $4 | 200K, text and images | Dominates A by the rule, but 4 points is inside the noise. Read it as “no reason to prefer A”. |
Row F is the one people misread. The label is correct, and the difference is not measurable.
What does “unrated” mean?
Unrated means no independent quality score exists for that model on the chosen test. It is not zero, and it is not a clean bill of health. On 23 September 2026, 210 of the 346 standard-delivery models in the catalogue (61%) had no score on the LMArena general board; the catalogue’s other 91 rows are batch and free listings. The board lists them by price and keeps them out of the quality order.
An unrated model can be excellent. You simply cannot use this board to prove it, in either direction.
What should you do with a dominated verdict?
A verdict is evidence for a decision, not the decision. The discipline I use for that step is ROUTE, from my book The Model Portfolio, which is available in English and Dutch. ROUTE is a five-step way to run a mix of AI models on purpose instead of by habit. You do not need the book to use it.
- Register. Keep a list of every model you run, with its owner, its cost class, the data it may see, and its retirement status.
- Objective typing. Write down what each task needs (reasoning depth, output format, speed, quality bar, privacy) before you choose a model.
- Utilize policy. Write the rule that sends each task to a model, including what happens when that model fails.
- Track. Measure cost, speed, and quality per task, not only the monthly bill.
- Evolve. Promote, demote, and retire models on a schedule, with a way back.
A hypothetical worked example. Your support team summarises customer tickets with model A, and the board says C dominates A with nothing to give up.
- Register: confirm that A is really what runs, on which endpoint, and who owns the task.
- Objective typing: the task reads tickets of up to 8,000 tokens, returns a 150-word summary, and may only process data in the EU. The board checks none of that last requirement.
- Utilize policy: keep A in place. Run C as a trial on a copy of last week’s tickets.
- Track: compare cost per summary and the share of summaries your team rejects.
- Evolve: if C holds up, switch at the next review, keep A as the fallback for a month, and record the decision with its date.
ROUTE has its own page, and What is model routing? covers the routing policy itself.
What the label does not tell you
It does not tell you how a model does on your task. The default lens is general human preference. The board also offers a document lens and per-task rankings, but none of them is your support queue.
It does not include your switching cost. The Check page asks for your monthly spend and switching cost only to estimate how many months a switch would take to pay back.
It does not last. A new model, a price change, or a different workload can flip a verdict at the next update. The index takes no cut of inference, no affiliate fees, and no paid placement (independence page), but its numbers are still a snapshot.
Try this today (10 minutes)
- Open Check a model and name the model you pay the most for.
- Set the workload closest to your real work.
- Write down one of four results: on the frontier, dominated with nothing lost, dominated with a named loss, or unrated.
- Add the date, the score gap, and one sentence on what you would have to test before switching.
That one line is worth more than a leaderboard screenshot, because it says what you checked and when. For the story of why the index exists, read Introducing Undominated.ai.
Cite this:What is a dominated AI model? Better and cheaper is only half the test.Len P. van der Hof. https://lenvanderhof.com/en/blog/what-is-a-dominated-model/ ·