Skip to content
fedi.software

fedi-index methodology

Sources, licences, weights and formulas of the index — rendered from the same configuration the calculation uses.

Methodology v1 · Live data for 2026-10 (recalculated daily) · History · JSON

  • Only sources with an explicit open licence (CC BY, MIT, Apache) enter the index; sources still waiting for permission stay switched off.
  • Independent runs count; vendor self-reports are shown with a label and get no weight.
  • One source family is one voice: correlated categories of the same voters do not add up.
  • Every score has a public decomposition: component, source, value, weight and date.
  • Uncertainty is visible: a score comes with its interval, little data means “provisional”.

Sources and licences

Source Family Licence In the index
LMArena Arena Leaderboard Dataset by LMArena, CC BY 4.0. Modified by fedi.software (equated, weighted, aggregated). Not endorsed by LMArena. Arena (human votes) CC BY 4.0 yes
Epoch AI Epoch AI, 'Capabilities & benchmarking'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks'. CC BY. ECI by Epoch AI; Epoch-run results only. Modified by fedi.software. Epoch AI (benchmarks) CC BY 4.0 yes
τ²-bench (Sierra) τ²-bench leaderboard, © Sierra Research, MIT License (Yao et al. 2024, arXiv:2406.12045; Barres et al. 2025, arXiv:2506.07982). Sierra-run text submissions only. τ²-bench (agents) MIT yes
LiveBench LiveBench — waiting for permission
Terminal-Bench Terminal-Bench — waiting for permission
SWE-rebench SWE-rebench — waiting for permission
Toolathlon Toolathlon — waiting for permission
models.dev Pricing: models.dev, MIT License, © 2025 models.dev. modelsdev MIT yes
LiteLLM Pricing: LiteLLM model_prices_and_context_window.json, MIT License, © 2023 Berri AI. litellm MIT yes
Hugging Face Downloads (30 days): Hugging Face Hub, as of the snapshot date. hf — yes

Slices and weights

General

Arena (human votes) share 50%

  • text/overall text/overall2
  • text/hard_prompts text/hard_prompts2
  • text/expert text/expert1
  • text/instruction_following text/instruction_following1
  • text/multi_turn text/multi_turn0.50
  • text/longer_query text/longer_query0.50

Epoch AI (benchmarks) share 50%

  • Epoch Capabilities Index eci3
  • GPQA Diamond gpqa_diamond0.50
  • SimpleQA Verified simpleqa_verified0.50

Coding

Arena (human votes) share 60%

  • text/coding text/coding1.5
  • webdev/overall webdev/overall1.5

Epoch AI (benchmarks) share 40%

  • SWE-bench Verified swe_bench_verified1

Agents

Arena (human votes) share 60%

  • agent/overall agent/overall1

τ²-bench (agents) share 40%

  • τ-Knowledge banking tau2/banking_knowledge1.5
  • τ²-bench retail tau2/retail1
  • τ²-bench airline tau2/airline1
  • τ²-bench telecom tau2/telecom1

Math

Arena (human votes) share 50%

  • text/math text/math1

Epoch AI (benchmarks) share 50%

  • FrontierMath Tiers 1–3 frontiermath_t1_32
  • FrontierMath Tier 4 frontiermath_t41

Vision provisional

Arena (human votes) share 100%

  • vision/overall vision/overall2
  • vision/ocr vision/ocr0.50
  • vision/diagram vision/diagram0.50

Search provisional

Arena (human votes) share 100%

  • search/overall search/overall1

Writing provisional

Arena (human votes) share 100%

  • text/creative_writing text/creative_writing2
  • text/longer_query text/longer_query1

Languages provisional

Arena votes in each language; a slice is shown on the site version in that language.

English — text/english · Русский — text/russian · Deutsch — text/german · Español — text/spanish · Français — text/french · 日本語 — text/japanese · 한국어 — text/korean · 繁體中文 — text/chinese · Polski — text/polish

Benchmarks and their age

Component Version released
Epoch Capabilities Index eci live (no decay)
GPQA Diamond gpqa_diamond 2023-11-20
SimpleQA Verified simpleqa_verified 2025-09-09
FrontierMath Tiers 1–3 frontiermath_t1_3 2026-06-12
FrontierMath Tier 4 frontiermath_t4 2026-06-12
SWE-bench Verified swe_bench_verified 2024-08-13
τ²-bench retail tau2/retail 2024-06-17
τ²-bench airline tau2/airline 2024-06-17
τ²-bench telecom tau2/telecom 2025-06-09
τ-Knowledge banking tau2/banking_knowledge 2026-03-04

An older benchmark version gets less weight (contamination, saturation): the weight halves every decay_half_life months, down to decay_min.

Formulas

Every component is moved to the Arena scale by linear equating on the models measured by both (anchor A = Arena text overall). The parameters are fixed on the reference snapshot of the methodology version, so months stay comparable.

x_eq  = μ_A + (x − μ_c) · σ_A / σ_c          se_eq = se · σ_A / σ_c
w_c   = 8 · family_share[f] · base_w[c] / Σ base_w(f, present)
rel_c = τ² / (τ² + se_eq²)
w_eff = w_c · rel_c · decay_c · (self_report ? 0 : 1)
Q     = (Σ w_eff · x_eq + w0 · prior) / (Σ w_eff + w0)
SE(Q) = √Σ (w_eff · se_eq)² / (Σ w_eff + w0)

Within a slice the components are averaged with their weights, reliability (τ) and the benchmark age; a model with little data is pulled towards the prior. Parameters:

anchor
arena:text/overall
total_weight
8
tau
15
w0
0.50
prior_pct
25
min_anchors
8
min_votes
300
min_confidence
0.90
decay_half_life
18
decay_min
0.25
outlier_sigma
3
outlier_factor
0.50
population_months
18
jump_sigma
3

Tier = the worse of the tier by percentile and the tier by the confidence interval; a model with one source family gets at most A.

Tier = the worse of the tier by percentile and the tier by the confidence interval; a model with one source family gets at most A.

provisional_max
A

S ≥ p95A ≥ p80B ≥ p50C ≥ p25 D < p25

Tier by the confidence interval (upper bound of the rank / models): S ≤ 10% · A ≤ 30% · B ≤ 60% · C ≤ 85%

Value: quality for the price

Value is the distance above the regression line of quality on log₂ of the price; the price is the median of providers.

slice
general
input_share
3
output_share
1
price_sources
modelsdev, litellm
unit
1m_tokens

On your hardware

Local: only open weights, the best quantization that fits a stack, the score scaled by generation speed.

slice
general
context
8,192
speed_ref
20
low_penalty
0.60
  • GPU 12 GB
  • GPU 24 GB
  • Mac 64 GB
  • 128 GB unified memory
  • CPU, 32 GB RAM

Data, licence and archive

The index (our numbers and decompositions) is CC BY 4.0: “fedi-index by fedi.software” with a link. The components keep the licences of their sources, listed in sources[] of the JSON.

A snapshot is frozen on the 1st of every month and never changes; a new methodology version starts a new series.

Monthly JSON: 2026-10

Changelog

  • v1 2026-10-01 First version: LMArena + Epoch AI + τ²-bench, linear equating to Arena text Elo with parameters fixed on the reference snapshot, family shares, sub-scores arena / bench, value = residual of the price regression + Pareto, local = Fit on 5 stacks.
  • v1.1 2026-10-01 Local quant floor: the quant is picked like local-fit (Q3 floor; below it only when nothing of Q3+ fits — "strong compression", ×0.6, ranked after every Q3+ model). Local only; the v1 scale and the published 2026-10 snapshot are unchanged.