跳到主要內容
fedi.software

fedi-index 方法論

指數的來源、授權、權重與公式——直接取自計算所用的同一份設定。

方法論 v1 · 2026-10 即時資料(每日重新計算) · 歷史 · JSON

  • 只有具明確開放授權(CC BY、MIT、Apache)的來源會納入指數;仍在等待許可的來源維持關閉。
  • 只計入獨立測試結果;廠商自報數據會加上標示,但不給予權重。
  • 一個來源系列只算一票:同一批投票者的相關類別不會重複累加。
  • 每個分數都有公開拆解:分項、來源、數值、權重與日期。
  • 不確定性清楚可見:分數附帶區間,資料少就標示「暫定」。

來源與授權

來源 系列 授權 納入指數
LMArena Arena Leaderboard Dataset by LMArena, CC BY 4.0. Modified by fedi.software (equated, weighted, aggregated). Not endorsed by LMArena. Arena(人類投票) CC BY 4.0 是
Epoch AI Epoch AI, 'Capabilities & benchmarking'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks'. CC BY. ECI by Epoch AI; Epoch-run results only. Modified by fedi.software. Epoch AI(基準測試) CC BY 4.0 是
τ²-bench (Sierra) τ²-bench leaderboard, © Sierra Research, MIT License (Yao et al. 2024, arXiv:2406.12045; Barres et al. 2025, arXiv:2506.07982). Sierra-run text submissions only. τ²-bench(代理) MIT 是
LiveBench LiveBench — 等待許可
Terminal-Bench Terminal-Bench — 等待許可
SWE-rebench SWE-rebench — 等待許可
Toolathlon Toolathlon — 等待許可
models.dev Pricing: models.dev, MIT License, © 2025 models.dev. modelsdev MIT 是
LiteLLM Pricing: LiteLLM model_prices_and_context_window.json, MIT License, © 2023 Berri AI. litellm MIT 是
Hugging Face Downloads (30 days): Hugging Face Hub, as of the snapshot date. hf — 是

分項與權重

綜合

Arena(人類投票) 占比 50%

  • text/overall text/overall2
  • text/hard_prompts text/hard_prompts2
  • text/expert text/expert1
  • text/instruction_following text/instruction_following1
  • text/multi_turn text/multi_turn0.50
  • text/longer_query text/longer_query0.50

Epoch AI(基準測試) 占比 50%

  • Epoch Capabilities Index eci3
  • GPQA Diamond gpqa_diamond0.50
  • SimpleQA Verified simpleqa_verified0.50

程式設計

Arena(人類投票) 占比 60%

  • text/coding text/coding1.5
  • webdev/overall webdev/overall1.5

Epoch AI(基準測試) 占比 40%

  • SWE-bench Verified swe_bench_verified1

代理

Arena(人類投票) 占比 60%

  • agent/overall agent/overall1

τ²-bench(代理) 占比 40%

  • τ-Knowledge banking tau2/banking_knowledge1.5
  • τ²-bench retail tau2/retail1
  • τ²-bench airline tau2/airline1
  • τ²-bench telecom tau2/telecom1

數學

Arena(人類投票) 占比 50%

  • text/math text/math1

Epoch AI(基準測試) 占比 50%

  • FrontierMath Tiers 1–3 frontiermath_t1_32
  • FrontierMath Tier 4 frontiermath_t41

視覺 暫定

Arena(人類投票) 占比 100%

  • vision/overall vision/overall2
  • vision/ocr vision/ocr0.50
  • vision/diagram vision/diagram0.50

搜尋 暫定

Arena(人類投票) 占比 100%

  • search/overall search/overall1

寫作 暫定

Arena(人類投票) 占比 100%

  • text/creative_writing text/creative_writing2
  • text/longer_query text/longer_query1

語言 暫定

各語言的 Arena 投票;該分項會顯示在對應語言版本的網站上。

English — text/english · Русский — text/russian · Deutsch — text/german · Español — text/spanish · Français — text/french · 日本語 — text/japanese · 한국어 — text/korean · 繁體中文 — text/chinese · Polski — text/polish

基準測試與其年齡

分項 版本發布
Epoch Capabilities Index eci 即時(不衰減)
GPQA Diamond gpqa_diamond 2023-11-20
SimpleQA Verified simpleqa_verified 2025-09-09
FrontierMath Tiers 1–3 frontiermath_t1_3 2026-06-12
FrontierMath Tier 4 frontiermath_t4 2026-06-12
SWE-bench Verified swe_bench_verified 2024-08-13
τ²-bench retail tau2/retail 2024-06-17
τ²-bench airline tau2/airline 2024-06-17
τ²-bench telecom tau2/telecom 2025-06-09
τ-Knowledge banking tau2/banking_knowledge 2026-03-04

較舊的基準測試版本權重較低(資料污染、飽和):權重每 decay_half_life 個月減半,最低為 decay_min。

公式

每個分項都透過線性等化換算到 Arena 尺度,依據是兩者都測量過的模型(錨點 A = Arena text overall)。參數固定在方法論版本的參考快照上,因此各月份可互相比較。

x_eq  = μ_A + (x − μ_c) · σ_A / σ_c          se_eq = se · σ_A / σ_c
w_c   = 8 · family_share[f] · base_w[c] / Σ base_w(f, present)
rel_c = τ² / (τ² + se_eq²)
w_eff = w_c · rel_c · decay_c · (self_report ? 0 : 1)
Q     = (Σ w_eff · x_eq + w0 · prior) / (Σ w_eff + w0)
SE(Q) = √Σ (w_eff · se_eq)² / (Σ w_eff + w0)

在分項內,各分項依權重、可靠度(τ)與基準測試年齡加權平均;資料少的模型會被拉向先驗值。參數:

anchor
arena:text/overall
total_weight
8
tau
15
w0
0.50
prior_pct
25
min_anchors
8
min_votes
300
min_confidence
0.90
decay_half_life
18
decay_min
0.25
outlier_sigma
3
outlier_factor
0.50
population_months
18
jump_sigma
3

等級 = 百分位等級與信賴區間等級中較差者;只有一個來源系列的模型最高為 A。

等級 = 百分位等級與信賴區間等級中較差者;只有一個來源系列的模型最高為 A。

provisional_max
A

S ≥ p95A ≥ p80B ≥ p50C ≥ p25 D < p25

依信賴區間的等級(排名上限 / 模型數): S ≤ 10% · A ≤ 30% · B ≤ 60% · C ≤ 85%

性價比:以價格衡量品質

性價比是品質對價格 log₂ 迴歸線之上的距離;價格為各供應商中位數。

slice
general
input_share
3
output_share
1
price_sources
modelsdev, litellm
unit
1m_tokens

在你的硬體上

本機:僅限開放權重,採用能載入該配置的最佳量化,分數依生成速度調整。

slice
general
context
8,192
speed_ref
20
low_penalty
0.60
  • GPU 12 GB
  • GPU 24 GB
  • Mac 64 GB
  • 128 GB 統一記憶體
  • CPU,32 GB RAM

資料、授權與封存

指數(我們的數字與拆解)採 CC BY 4.0:「fedi-index by fedi.software」並附連結。各分項沿用其來源的授權,列於 JSON 的 sources[]。

每月 1 日凍結一次快照,之後不再變更;新的方法論版本會開啟新的系列。

每月 JSON: 2026-10

變更記錄

  • v1 2026-10-01 First version: LMArena + Epoch AI + τ²-bench, linear equating to Arena text Elo with parameters fixed on the reference snapshot, family shares, sub-scores arena / bench, value = residual of the price regression + Pareto, local = Fit on 5 stacks.
  • v1.1 2026-10-01 Local quant floor: the quant is picked like local-fit (Q3 floor; below it only when nothing of Q3+ fits — "strong compression", ×0.6, ranked after every Q3+ model). Local only; the v1 scale and the published 2026-10 snapshot are unchanged.