Skip to content
fedi.software

Agents

Models ranked on agentic tasks — tool calls and multi-step work — by the LMArena agent leaderboard. Its score is centred on zero, not an Elo.

Updated · Sources: LMArena, LiteLLM, models.dev

Cost per month
Tier Model Score Input / output, per 1M Per month Context App Compare
S0.143$10 / $50 1M
S0.138$4 / $20 1M
S0.125$2 / $10 1M
S0.087$5 / $25 1M
S0.082$3 / $18.50 1M
A0.066$0.425 / $2.12 1M
A0.044$1.44 / $7.20 1M
A0.040$0.10 / $0.20 1M
A0.035$0.30 / $1.05Free 1M
B0.024$0.40 / $1.40Free 1M
B-0.004$0.03 / $0.10Free 1M
C-0.033$0.10 / $0.20 1M
C-0.049$1.25 / $4.25 1M
Compare ()

Tiers are ours: a model's place within the board counting only the models whose whole confidence interval is above it. S — top 10%, A — up to 30%, B — up to 60%, C — up to 85%, D — the rest.

Quality vs price rating up, blended price (3 parts input : 1 output) across, USD per 1M tokens
-0.2 -0.1 0.0 0.1 0.2 $0.01 $0.10 $1 $10 Claude Fable 5.1 Claude Opus 5.5 Claude Sonnet 5.5

What does the agents leaderboard rank?

It follows the LMArena agent leaderboard: models are given tasks that take several steps and require calling tools — searching, reading files, running code — and people judge which model handled the task better. Instead of single answers, whole sessions are compared.

Why is there a Score instead of an Elo?

This leaderboard publishes a score centred on zero rather than an Elo rating. A model above zero was preferred more often than the average model of the board, a model below zero less often. The column is labelled "Score" and shows three decimals, and the votes counted are sessions. Because the scale is different, never compare these numbers with the Elo ratings of the text or coding boards; compare models within this board only.

Are tiers calculated the same way?

Yes. The tier rule depends only on the confidence intervals and the share of the board, not on the scale. A model needs 300 sessions to be tiered. Its place is one plus the number of models whose whole interval lies above its own; S goes to the first 10% of places, A up to 30%, B up to 60%, C up to 85%, D to the rest.

What else does each row show?

  • the vendor and whether the weights are open;
  • the cheapest price per million tokens among providers, from models.dev and LiteLLM;
  • the context window — long contexts matter for agents that read a lot;
  • the price trend over 30 days.

What should I be careful about?

The board is younger than the text arena and lists fewer models. The agent framework used in the arena is also not the one you will use: a model that performs well there can behave differently inside your own coding agent or automation. Check the coding agents page for the tools themselves, and try the shortlisted models on a real task before committing.