Skip to content
fedi.software

fedi-index: Agents

Models as agents: tool use and multi-step tasks — Arena agent sessions and τ²-bench runs.

Live data for 2026-10 (recalculated daily) · Methodology v1 · Methodology · History · JSON

All Open weights Closed
# Tier Model fedi-index Percentile People prefer Solves tasks Value Sources Score decomposition
13B1,456 ±5791,4591,449 +44 Aτ›
14B1,454 ±5771,4561,449 +45 Aτ›
17B1,449 ±2721,459— +92◆ A›
19B1,448 ±3681,458— — A›
24B1,442 ±3601,451— +60 A›
27B1,435 ±7541,445— +33 A›
30C1,430 ±2491,439— +105◆ A›
32C1,428 ±3461,437— — A›
33C1,428 ±6441,438— +98◆ A›
36C1,424 ±3391,433— +46 A›
42C1,408 ±5281,416— +68 A›
47D1,402 ±4191,410— +50 A›
49D1,399 ±4161,407— +80 A›
53D1,388 ±87—1,411 +23 τ›
54D1,388 ±87—1,411 +27 τ›
55D1,386 ±651,394— +38 A›
56D1,384 ±441,392— +27 A›
57D1,377 ±621,384— -19 A›

…

AArena (human votes)ττ²-bench (agents) ◆ Pareto frontier: nothing is both better and cheaper

fedi-index is an Elo-equivalent on a fixed scale: comparable between months within one methodology version. Tiers S–D: by percentile in the slice, and no better than the confidence interval allows (cut-offs in the methodology). “Provisional” — only one source family measured the model here. Methodology

Embed on your site

The table as a widget for your site: it updates by itself and keeps the attribution of the sources.

JSON fedi-index © fedi.software, CC BY 4.0; components keep the licences of their sources.