Skip to content
fedi.software

Text

Chat models ranked by blind human votes on LMArena, with the confidence interval of every rating and the current price per million tokens.

Updated · Sources: LMArena, LiteLLM, models.dev

Cost per month
Tier Model Elo Input / output, per 1M Per month Context App Compare
S1,471$0.40 / $1.40Free 1M
S1,470$0.30 / $1.05Free 1M
S1,470$0.03 / $0.10Free 1M
S1,461$0.45 / $2.15 200K
S1,451$1.25 / $2.50 1M
A1,450$1.25 / $2.50 1M
A1,448$2 / $6 500K
A1,446$0.40 / $1.75 205K
A1,444$1.25 / $2.50 2M
A1,440$0.2740 / $1.10 205K
A1,437$0.7042 / $3.10 200K
A1,437$2 / $10 200K
A1,436$0.2740 / $1.10 205K
A1,430$0.286 / $1.14 131K
A1,427$1.25 / $6 500K
A1,426$3 / $15 131K
A1,411$2.70 / $13.50 256K
B1,400$1.60 / $4.80 500K
B1,397$1.25 / $2.50 1M
B1,384$0.1165 / $0.2911 131K
B1,377$0.137 / $0.411 128K
B1,369$0.30 / $0.50 131K
B1,351$0.06 / $0.40 200K
B1,333$0.29 / $0.86 64K
B1,331$10.00 / $10.00 128K
B1,305$2 / $10 131K
B1,281— —
C1,226— —
D972— —
D939— —
D918— —
Compare ()

Tiers are ours: a model's place within the board counting only the models whose whole confidence interval is above it. S — top 10%, A — up to 30%, B — up to 60%, C — up to 85%, D — the rest.

Quality vs price rating up, blended price (3 parts input : 1 output) across, USD per 1M tokens
1,000 1,200 1,400 1,600 $0.01 $0.10 $1 $10 $100 Claude Opus 5.5 Claude Fable 5.1 Claude Opus 5

What the text leaderboard shows

This page ranks chat models by the LMArena text arena, overall category. The method is simple to describe: a visitor sends a prompt, two anonymous models answer, the visitor picks the better reply, and only then are the names revealed. Hundreds of thousands of such votes are turned into an Elo-style rating, the same idea that ranks chess players.

Each row combines that rating with facts from other sources: the vendor, a mark for open weights, the cheapest current price per million input and output tokens we found among providers, the context window, the price trend over the last 30 days and, when the vendor has a desktop or phone app in our catalogue, a link to download it.

How to read a row

  • Tier. Our letter from S to D, explained below.
  • Elo and its bar. The bar shows the 95% confidence interval. When the intervals of two models overlap, the difference between them is not statistically clear.
  • Price. The lowest price among providers on the day of the data, in USD, for input and for output.
  • Context. How many tokens the model accepts in one request.

A click on a row opens the details: providers sorted by price, modalities, maximum output, the number of votes, the rank at the source, licence, release date and knowledge cutoff.

How tiers are assigned

Only models with at least 300 votes take part. For each model we count how many others have their whole confidence interval above its interval; that count, divided by the number of models, gives its place. Under 10% means S, under 30% A, under 60% B, under 85% C, and the rest D. Models statistically tied with the leader therefore share the top tier. A newer model with fewer votes is shown as "Not tiered yet".

Tools on the page

The chips filter by open weights, free access, low price (a blended price under 1 USD per million tokens) and vendor. The calculator turns your monthly token volume and the share of input into a cost per month for every model, and the chart plots quality against a blended price of 3 parts input to 1 part output. Tick up to 4 models to compare them in detail.

What the ranking does not tell you

Arena votes measure which answer people liked, mostly on English prompts. They say nothing about speed, and little about a narrow task such as legal drafting or maths. Prices at a provider can also differ from the vendor's own list: caching, batch discounts and regional prices are not shown.