Skip to content
fedi.software

Coding

Models ranked on programming prompts: tier, Elo with its confidence interval, price per million tokens and context window.

Updated · Sources: LMArena, LiteLLM, models.dev

Cost per month
Tier Model Elo Input / output, per 1M Per month Context App Compare
S1,510$1 / $4Free 1M
S1,489$0.275 / $1.10Free 262K
S1,483— —
S1,483$0.30 / $1.90 262K
A1,453$1.15 / $8 262K
A1,433— —
B1,399$0.40 / $1.80 256K
B1,386$0.30 / $2.50 1M
B1,381— —
B1,375— —
B1,312— —
B1,312— —
C1,270— —
C1,267— —
C1,239— —
C1,230— —
C1,218— —
C1,190— —
C1,169— —
C1,165— —
C1,159— —
C1,154— —
C1,143— —
C1,132— —
C1,129— —
C1,116— —
C1,113— —
C1,112— —
C1,098— —
C1,089— —
C1,080— —
D1,078— —
D1,067— —
D1,064— —
–1,053— —
D1,048— —
D1,045— —
D1,041— —
–1,039— —
D1,010— —
D1,001— —
–996— —
D945— —
D928— —
D913— —
D900— —
D886— —
D799— —
D776— —
D764— —
Compare ()

Tiers are ours: a model's place within the board counting only the models whose whole confidence interval is above it. S — top 10%, A — up to 30%, B — up to 60%, C — up to 85%, D — the rest.

Quality vs price rating up, blended price (3 parts input : 1 output) across, USD per 1M tokens
1,000 1,200 1,400 1,600 $0.01 $0.10 $1 $10 $100 MiMo-V2.6-Pro Claude Opus 5.5 Claude Opus 4.6

What is this leaderboard?

It ranks models by the coding category of the LMArena text arena. The voting works as on the general text board — two anonymous answers, one human vote — but only prompts about programming count: writing a function, explaining an error, refactoring a snippet. The result is an Elo rating with a 95% confidence interval for every model.

How is it different from the web dev and agents boards?

This board rates answers in a chat about code. Building a whole working web app from a prompt is measured separately on the web dev board, and multi-step work with tool calls on the agents board. A model can do well in one and less well in another, so look at all three if you are picking a model for a coding assistant.

What do the columns mean?

  • the tier letter, from S to D;
  • the Elo rating with its confidence bar;
  • the cheapest price per million input and output tokens among providers;
  • the context window, which matters when you feed a model large parts of a repository;
  • the price trend over 30 days.

The cost calculator is useful here: coding sessions send much more input than they receive, so set the input share high and see what a month of use would cost with each model.

How are tiers calculated?

With the same rule as on every leaderboard of the section. A model needs at least 300 votes. Its place is one plus the number of models whose entire confidence interval lies above its own. Divided by the total, a place under 10% gives S, under 30% A, under 60% B, under 85% C, otherwise D. Tied leaders share S.

Where do the prices come from?

From models.dev and LiteLLM, refreshed daily. We show the lowest current offer for each model and keep the history, which is how the trend column is drawn.

What should I keep in mind?

Voting on chat answers is not the same as running tests on a real project. A model that writes a neat explanation may still produce code that fails in your codebase, and the arena prompts are short compared with a day of real work. Use the board to narrow the list, then try two or three models on your own code.