AI Systems Research guide

Introducing Undominated.ai: the independent AI inference price index

Best first. Then price. The board that names the models that are both worse and dearer.

Warehouse floor, eleven crates in a chalk frontier under daylight, a large unused pile in shadow
Eleven on the edge. The rest of the pile is both worse and dearer.

You already have a model in production. Someone on the team found a cheaper endpoint. The spreadsheet treats that as a win. It is not a win if the cheaper row cannot do the job, and it is not a win if a third row is both stronger and cheaper than the one you kept.

Undominated.ai is the index I built for that decision. Capability first. Price second. The row that is both worse and dearer gets named.

What it is

The public line is short on purpose: Undominated.ai: the AI price index. Best first. Then price.

It is not a router. It does not sell tokens. It does not take a percentage of inference. It does not put a paid badge on a row. The independence page is the commercial boundary in one screen. Revenue in v1 is none. A later paid feed, if it exists, would sell history and provenance, not a better rank.

Capability comes from independent evaluators. Two sources sit side by side and are never blended: Artificial Analysis indices, shown for evaluation, and LMArena human preference, used under CC BY 4.0 from the official dataset. The site does not score models itself. Where the two sources part company, that disagreement is the finding. A composite would erase it.

Why an average is the wrong object

The spread between the cheapest and the dearest model in the catalogue is about 15,000×. That is not a rounding error. It is why a single “value” number is a marketing instrument. Blend quality and price into one score and the worse-and-dearer row can hide behind an attractive ratio. Quality-per-dollar, run on this catalogue, crowns cheap models near the floor of the scale. Nobody should ship those as “best value.”

Other public indexes exist. Some publish a geometric-mean blended $/M. Some chain SKUs like a commodity index. Some weight lab tokens by transacted volume. They are answering a different question. This one asks which models are still worth considering at all, then what the job costs under a workload you can see.

What “undominated” means

A model is dominated when another model in the catalogue scores higher and costs less and can still do what it can do: same context class, same modalities, same tools. Raw two-axis Pareto lies often. The methodology reserves “better and cheaper” for a replacement whose capability envelope is a superset. If you would give something up, the board says so.

A model is on the value frontier when nothing both beats it and undercuts it under the selected lens and workload. Chartreuse in the interface is reserved for that membership. It is not a highlight colour for sale.

Ranks are significance ranks, the way LMArena’s Rank (UB) has worked for years. Models the measurement cannot separate share a rank. “The #1 model” is not a claim this data supports. Ties are ties.

Unrated is not zero. Most of the catalogue has no independent quality score. Those rows are listed by price and kept out of quality order. Not measured is not the same as measured badly.

A snapshot, not a monument

On 25 August 2026 the live board reported:

  • 409 models, 50 providers, 55 with tiered pricing
  • 97 of 108 rated, priced models beaten on quality and undercut on price
  • 11 not. That set is the frontier.

Those numbers move. Workload shape moves them. Batch delivery moves them. A new fetch moves them. Cite the live page, not this paragraph.

Headline $/M is often the wrong number. Context tiers change the rate past a prompt length. Some models bill reasoning tokens on a separate meter. A few discount by time of day. The board reprices the row when the workload you selected crosses those lines.

How to use it this week

  1. Open the leaderboard. Set the lens and the workload to the job you actually run.
  2. Open the frontier. That is the short list.
  3. Check the model you already pay for. If it is dominated, the board names what beats it and what you would give up.
  4. Read the methodology before you treat a rank as a purchase order.
  5. If a live figure disagrees with a primary source, file a correction. The public GitHub repo is the tracker, not a second copy of the catalogue.

English is the source language. Locales ship as machine translation with numbers, model names, and provider names left untouched.

If a cheaper endpoint cannot do the job, it was never cheaper. If something else is both stronger and cheaper, keeping the incumbent is a preference, not a cost argument. Put the preference on the table. Then look at the eleven.

Sources

  1. Undominated.ai
  2. Methodology
  3. Independence
  4. Public correction tracker
  5. LMArena ranking method (Rank UB) · LMArena

Further reading

Markdown for LLMs