本文へスキップ
fedi.software

fedi-indexの算出方法

インデックスのソース、ライセンス、重み、計算式 — 実際の計算と同じ設定から表示しています。

算出方法 v1 · 2026-10のライブデータ(毎日再計算) · 推移 · JSON

  • インデックスに含めるのは明示的なオープンライセンス(CC BY、MIT、Apache)を持つソースのみです。許可待ちのソースは無効のままです。
  • 独立した実行結果のみを採用します。ベンダーの自己申告はラベル付きで表示し、重みは付けません。
  • 1つのソースファミリーは1票:同じ投票者による相関したカテゴリは合算しません。
  • すべてのスコアに公開された内訳があります:要素、ソース、値、重み、日付。
  • 不確実性を可視化:スコアには区間が付き、データが少なければ「暫定」と表示します。

ソースとライセンス

ソース ファミリー ライセンス インデックスに採用
LMArena Arena Leaderboard Dataset by LMArena, CC BY 4.0. Modified by fedi.software (equated, weighted, aggregated). Not endorsed by LMArena. Arena(人間の投票) CC BY 4.0 はい
Epoch AI Epoch AI, 'Capabilities & benchmarking'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks'. CC BY. ECI by Epoch AI; Epoch-run results only. Modified by fedi.software. Epoch AI(ベンチマーク) CC BY 4.0 はい
τ²-bench (Sierra) τ²-bench leaderboard, © Sierra Research, MIT License (Yao et al. 2024, arXiv:2406.12045; Barres et al. 2025, arXiv:2506.07982). Sierra-run text submissions only. τ²-bench(エージェント) MIT はい
LiveBench LiveBench — 許可待ち
Terminal-Bench Terminal-Bench — 許可待ち
SWE-rebench SWE-rebench — 許可待ち
Toolathlon Toolathlon — 許可待ち
models.dev Pricing: models.dev, MIT License, © 2025 models.dev. modelsdev MIT はい
LiteLLM Pricing: LiteLLM model_prices_and_context_window.json, MIT License, © 2023 Berri AI. litellm MIT はい
Hugging Face Downloads (30 days): Hugging Face Hub, as of the snapshot date. hf — はい

カテゴリと重み

総合

Arena(人間の投票) 割合 50%

  • text/overall text/overall2
  • text/hard_prompts text/hard_prompts2
  • text/expert text/expert1
  • text/instruction_following text/instruction_following1
  • text/multi_turn text/multi_turn0.50
  • text/longer_query text/longer_query0.50

Epoch AI(ベンチマーク) 割合 50%

  • Epoch Capabilities Index eci3
  • GPQA Diamond gpqa_diamond0.50
  • SimpleQA Verified simpleqa_verified0.50

コーディング

Arena(人間の投票) 割合 60%

  • text/coding text/coding1.5
  • webdev/overall webdev/overall1.5

Epoch AI(ベンチマーク) 割合 40%

  • SWE-bench Verified swe_bench_verified1

エージェント

Arena(人間の投票) 割合 60%

  • agent/overall agent/overall1

τ²-bench(エージェント) 割合 40%

  • τ-Knowledge banking tau2/banking_knowledge1.5
  • τ²-bench retail tau2/retail1
  • τ²-bench airline tau2/airline1
  • τ²-bench telecom tau2/telecom1

数学

Arena(人間の投票) 割合 50%

  • text/math text/math1

Epoch AI(ベンチマーク) 割合 50%

  • FrontierMath Tiers 1–3 frontiermath_t1_32
  • FrontierMath Tier 4 frontiermath_t41

ビジョン 暫定

Arena(人間の投票) 割合 100%

  • vision/overall vision/overall2
  • vision/ocr vision/ocr0.50
  • vision/diagram vision/diagram0.50

検索 暫定

Arena(人間の投票) 割合 100%

  • search/overall search/overall1

文章作成 暫定

Arena(人間の投票) 割合 100%

  • text/creative_writing text/creative_writing2
  • text/longer_query text/longer_query1

言語 暫定

各言語でのArena投票。カテゴリはその言語版のサイトに表示されます。

English — text/english · Русский — text/russian · Deutsch — text/german · Español — text/spanish · Français — text/french · 日本語 — text/japanese · 한국어 — text/korean · 繁體中文 — text/chinese · Polski — text/polish

ベンチマークとその古さ

要素 バージョン公開日
Epoch Capabilities Index eci ライブ(減衰なし)
GPQA Diamond gpqa_diamond 2023-11-20
SimpleQA Verified simpleqa_verified 2025-09-09
FrontierMath Tiers 1–3 frontiermath_t1_3 2026-06-12
FrontierMath Tier 4 frontiermath_t4 2026-06-12
SWE-bench Verified swe_bench_verified 2024-08-13
τ²-bench retail tau2/retail 2024-06-17
τ²-bench airline tau2/airline 2024-06-17
τ²-bench telecom tau2/telecom 2025-06-09
τ-Knowledge banking tau2/banking_knowledge 2026-03-04

古いベンチマークバージョンほど重みを下げます(学習データへの混入、飽和)。重みはdecay_half_lifeか月ごとに半減し、下限はdecay_minです。

計算式

各要素は、両方で測定されたモデルを使った線形等化でArenaのスケールに変換します(アンカーA = Arena text overall)。パラメータは算出方法バージョンの基準スナップショットで固定されるため、月をまたいで比較できます。

x_eq  = μ_A + (x − μ_c) · σ_A / σ_c          se_eq = se · σ_A / σ_c
w_c   = 8 · family_share[f] · base_w[c] / Σ base_w(f, present)
rel_c = τ² / (τ² + se_eq²)
w_eff = w_c · rel_c · decay_c · (self_report ? 0 : 1)
Q     = (Σ w_eff · x_eq + w0 · prior) / (Σ w_eff + w0)
SE(Q) = √Σ (w_eff · se_eq)² / (Σ w_eff + w0)

カテゴリ内では、各要素を重み・信頼性(τ)・ベンチマークの古さに応じて平均します。データの少ないモデルは事前値に近づけます。パラメータ:

anchor
arena:text/overall
total_weight
8
tau
15
w0
0.50
prior_pct
25
min_anchors
8
min_votes
300
min_confidence
0.90
decay_half_life
18
decay_min
0.25
outlier_sigma
3
outlier_factor
0.50
population_months
18
jump_sigma
3

ティア = パーセンタイルによるティアと信頼区間によるティアのうち低い方。ソースファミリーが1つのみのモデルは最高でAです。

ティア = パーセンタイルによるティアと信頼区間によるティアのうち低い方。ソースファミリーが1つのみのモデルは最高でAです。

provisional_max
A

S ≥ p95A ≥ p80B ≥ p50C ≥ p25 D < p25

信頼区間によるティア(順位の上限 / モデル数): S ≤ 10% · A ≤ 30% · B ≤ 60% · C ≤ 85%

コスパ:価格あたりの品質

コスパは、価格のlog₂に対する品質の回帰直線からどれだけ上にあるかです。価格はプロバイダーの中央値です。

slice
general
input_share
3
output_share
1
price_sources
modelsdev, litellm
unit
1m_tokens

手元のハードウェアで

ローカル:オープンウェイトのみ、各環境に収まる最良の量子化を使い、スコアを生成速度で補正します。

slice
general
context
8,192
speed_ref
20
low_penalty
0.60
  • GPU 12 GB
  • GPU 24 GB
  • Mac 64 GB
  • ユニファイドメモリ 128 GB
  • CPU、RAM 32 GB

データ、ライセンス、アーカイブ

インデックス(当サイトの数値と内訳)はCC BY 4.0です:「fedi-index by fedi.software」とリンクを表記してください。各要素はそれぞれのソースのライセンスに従い、JSONのsources[]に記載されています。

スナップショットは毎月1日に確定し、以後変更されません。新しい算出方法バージョンでは新しい系列が始まります。

月次JSON: 2026-10

変更履歴

  • v1 2026-10-01 First version: LMArena + Epoch AI + τ²-bench, linear equating to Arena text Elo with parameters fixed on the reference snapshot, family shares, sub-scores arena / bench, value = residual of the price regression + Pareto, local = Fit on 5 stacks.
  • v1.1 2026-10-01 Local quant floor: the quant is picked like local-fit (Q3 floor; below it only when nothing of Q3+ fits — "strong compression", ×0.6, ranked after every Q3+ model). Local only; the v1 scale and the published 2026-10 snapshot are unchanged.