Skip to content
fedi.software

Web dev

Models ranked by building working web apps from a prompt in LMArena's WebDev arena, with prices and context.

Updated · Sources: LMArena, LiteLLM, models.dev

Cost per month
Tier Model Elo Input / output, per 1M Per month Context App Compare
S1,671$0.338 / $1.01 1M
S1,637$0.10 / $0.20 262K
S1,620$0.0015 / $0.09Free 1M
A1,591$0.045 / $0.32 262K
A1,583$0.2088 / $0.4176 1M
A1,581$0.035 / $0.07 1M
A1,570$0.959 / $2.74 1M
B1,515$0.825 / $2.48 1M
B1,482$1.03 / $6.16 262K
B1,461$0.276 / $1.65 1M
B1,399$0.1644 / $0.9864Free 262K
C1,362$0.18 / $0.35 164K
C1,360$0.1126 / $0.9008Free 262K
C1,357$0.0846 / $0.6768 262K
C1,274$0.13 / $0.50Free 262K
C1,272$0.22 / $0.33 164K
C1,253$0.0564 / $0.4512 262K
C1,243$0.0282 / $0.282 1M
Compare ()

Tiers are ours: a model's place within the board counting only the models whose whole confidence interval is above it. S — top 10%, A — up to 30%, B — up to 60%, C — up to 85%, D — the rest.

Quality vs price rating up, blended price (3 parts input : 1 output) across, USD per 1M tokens
1,000 1,200 1,400 1,600 1,800 2,000 $0.01 $0.10 $1 $10 Claude Opus 5.5 GPT-6 Astra Claude Sonnet 5.5

What the web dev leaderboard measures

This page follows the LMArena WebDev arena, where models are asked to build a working web application from a single prompt — a to-do list, a landing page, a small game. Two anonymous models produce two apps, and the voter chooses the one that works and looks better. The votes become an Elo rating with a 95% confidence interval, as on the other arenas.

The WebDev arena has its own set of models. Some models that appear on the text or coding boards are missing here, and a few are rated only here, so the list does not simply repeat the coding board.

How to read it

The rating reflects the whole result as a person sees it: whether the app runs, whether it matches the request, how the layout looks. It does not judge code style on its own. That makes the board a good signal for tools that generate prototypes or front-end code, and a weaker one for back-end logic or large refactoring.

  • The tier letter summarises the position within this board.
  • The bar around the rating is the confidence interval; overlapping bars mean no clear difference.
  • Prices per million tokens, context window and the 30-day trend come from models.dev and LiteLLM.

Tiers

A model needs 300 votes to be tiered. Its place is counted from the models whose entire interval sits above its own, and the share of the board decides the letter: S under 10%, A under 30%, B under 60%, C under 85%, D for the rest.

Limits of the method

Arena apps are small, single-page projects built in one go. How a model behaves over a long session in a large codebase, or with your framework of choice, is outside what the vote can capture. Compare the result with the coding board and, if the model will work as an agent, with the agents board.