fedi-index 방법론
지수의 소스, 라이선스, 가중치와 공식 — 실제 계산에 쓰이는 것과 같은 설정에서 생성됩니다.
방법론 v1 · 2026-10 실시간 데이터(매일 재계산) · 추이 · JSON
- 명시적인 오픈 라이선스(CC BY, MIT, Apache)가 있는 소스만 지수에 반영합니다. 허가를 기다리는 소스는 꺼 둡니다.
- 독립적인 실행 결과만 반영합니다. 제조사 자체 발표는 표시와 함께 보여 주되 가중치를 주지 않습니다.
- 소스 계열 하나는 한 표: 같은 투표자의 상관된 카테고리는 합산하지 않습니다.
- 모든 점수의 구성이 공개됩니다: 구성 요소, 소스, 값, 가중치, 날짜.
- 불확실성을 드러냅니다: 점수에는 구간이 붙고, 데이터가 적으면 ‘잠정’으로 표시합니다.
소스와 라이선스
| 소스 | 계열 | 라이선스 | 지수 반영 |
|---|---|---|---|
| LMArena Arena Leaderboard Dataset by LMArena, CC BY 4.0. Modified by fedi.software (equated, weighted, aggregated). Not endorsed by LMArena. | Arena(사람 투표) | CC BY 4.0 | 예 |
| Epoch AI Epoch AI, 'Capabilities & benchmarking'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks'. CC BY. ECI by Epoch AI; Epoch-run results only. Modified by fedi.software. | Epoch AI(벤치마크) | CC BY 4.0 | 예 |
| τ²-bench (Sierra) τ²-bench leaderboard, © Sierra Research, MIT License (Yao et al. 2024, arXiv:2406.12045; Barres et al. 2025, arXiv:2506.07982). Sierra-run text submissions only. | τ²-bench(에이전트) | MIT | 예 |
| LiveBench | LiveBench | — | 허가 대기 중 |
| Terminal-Bench | Terminal-Bench | — | 허가 대기 중 |
| SWE-rebench | SWE-rebench | — | 허가 대기 중 |
| Toolathlon | Toolathlon | — | 허가 대기 중 |
| models.dev Pricing: models.dev, MIT License, © 2025 models.dev. | modelsdev | MIT | 예 |
| LiteLLM Pricing: LiteLLM model_prices_and_context_window.json, MIT License, © 2023 Berri AI. | litellm | MIT | 예 |
| Hugging Face Downloads (30 days): Hugging Face Hub, as of the snapshot date. | hf | — | 예 |
분야와 가중치
종합
Arena(사람 투표) 비중 50%
- text/overall
text/overall2 - text/hard_prompts
text/hard_prompts2 - text/expert
text/expert1 - text/instruction_following
text/instruction_following1 - text/multi_turn
text/multi_turn0.50 - text/longer_query
text/longer_query0.50
Epoch AI(벤치마크) 비중 50%
- Epoch Capabilities Index
eci3 - GPQA Diamond
gpqa_diamond0.50 - SimpleQA Verified
simpleqa_verified0.50
코딩
Arena(사람 투표) 비중 60%
- text/coding
text/coding1.5 - webdev/overall
webdev/overall1.5
Epoch AI(벤치마크) 비중 40%
- SWE-bench Verified
swe_bench_verified1
에이전트
Arena(사람 투표) 비중 60%
- agent/overall
agent/overall1
τ²-bench(에이전트) 비중 40%
- τ-Knowledge banking
tau2/banking_knowledge1.5 - τ²-bench retail
tau2/retail1 - τ²-bench airline
tau2/airline1 - τ²-bench telecom
tau2/telecom1
수학
Arena(사람 투표) 비중 50%
- text/math
text/math1
Epoch AI(벤치마크) 비중 50%
- FrontierMath Tiers 1–3
frontiermath_t1_32 - FrontierMath Tier 4
frontiermath_t41
비전 잠정
Arena(사람 투표) 비중 100%
- vision/overall
vision/overall2 - vision/ocr
vision/ocr0.50 - vision/diagram
vision/diagram0.50
검색 잠정
Arena(사람 투표) 비중 100%
- search/overall
search/overall1
글쓰기 잠정
Arena(사람 투표) 비중 100%
- text/creative_writing
text/creative_writing2 - text/longer_query
text/longer_query1
언어 잠정
언어별 Arena 투표; 해당 분야는 그 언어 버전 사이트에 표시됩니다.
English — text/english · Русский — text/russian · Deutsch — text/german · Español — text/spanish · Français — text/french · 日本語 — text/japanese · 한국어 — text/korean · 繁體中文 — text/chinese · Polski — text/polish
벤치마크와 경과 기간
| 구성 요소 | 버전 출시 |
|---|---|
Epoch Capabilities Index eci |
실시간(감쇠 없음) |
GPQA Diamond gpqa_diamond |
2023-11-20 |
SimpleQA Verified simpleqa_verified |
2025-09-09 |
FrontierMath Tiers 1–3 frontiermath_t1_3 |
2026-06-12 |
FrontierMath Tier 4 frontiermath_t4 |
2026-06-12 |
SWE-bench Verified swe_bench_verified |
2024-08-13 |
τ²-bench retail tau2/retail |
2024-06-17 |
τ²-bench airline tau2/airline |
2024-06-17 |
τ²-bench telecom tau2/telecom |
2025-06-09 |
τ-Knowledge banking tau2/banking_knowledge |
2026-03-04 |
오래된 벤치마크 버전일수록 가중치가 낮아집니다(학습 데이터 오염, 포화). 가중치는 decay_half_life개월마다 절반이 되며 decay_min 아래로는 내려가지 않습니다.
공식
각 구성 요소는 양쪽에서 모두 측정된 모델을 기준으로 선형 동등화를 거쳐 Arena 척도로 옮깁니다(앵커 A = Arena text overall). 매개변수는 방법론 버전의 기준 스냅샷에 고정되므로 월별 비교가 가능합니다.
x_eq = μ_A + (x − μ_c) · σ_A / σ_c se_eq = se · σ_A / σ_c w_c = 8 · family_share[f] · base_w[c] / Σ base_w(f, present) rel_c = τ² / (τ² + se_eq²) w_eff = w_c · rel_c · decay_c · (self_report ? 0 : 1) Q = (Σ w_eff · x_eq + w0 · prior) / (Σ w_eff + w0) SE(Q) = √Σ (w_eff · se_eq)² / (Σ w_eff + w0)
분야 안에서 구성 요소는 가중치, 신뢰도(τ), 벤치마크의 경과 기간에 따라 평균합니다. 데이터가 적은 모델은 사전값 쪽으로 당겨집니다. 매개변수:
anchor- arena:text/overall
total_weight- 8
tau- 15
w0- 0.50
prior_pct- 25
min_anchors- 8
min_votes- 300
min_confidence- 0.90
decay_half_life- 18
decay_min- 0.25
outlier_sigma- 3
outlier_factor- 0.50
population_months- 18
jump_sigma- 3
등급 = 백분위 기준 등급과 신뢰 구간 기준 등급 중 낮은 쪽; 소스 계열이 하나뿐인 모델은 최고 A입니다.
등급 = 백분위 기준 등급과 신뢰 구간 기준 등급 중 낮은 쪽; 소스 계열이 하나뿐인 모델은 최고 A입니다.
provisional_max- A
S ≥ p95A ≥ p80B ≥ p50C ≥ p25 D < p25
신뢰 구간 기준 등급(순위 상한 / 모델 수): S ≤ 10% · A ≤ 30% · B ≤ 60% · C ≤ 85%
가성비: 가격 대비 품질
가성비는 가격의 log₂에 대한 품질 회귀선보다 얼마나 위에 있는지입니다. 가격은 제공업체 중앙값입니다.
slice- general
input_share- 3
output_share- 1
price_sources- modelsdev, litellm
unit- 1m_tokens
내 하드웨어에서
로컬: 오픈 웨이트만, 각 환경에 적재되는 최상의 양자화를 쓰고 점수를 생성 속도로 보정합니다.
slice- general
context- 8,192
speed_ref- 20
low_penalty- 0.60
- GPU 12 GB
- GPU 24 GB
- Mac 64 GB
- 통합 메모리 128 GB
- CPU, RAM 32 GB
데이터, 라이선스, 아카이브
지수(자체 수치와 구성 내역)는 CC BY 4.0입니다: “fedi-index by fedi.software”와 링크를 표기하세요. 구성 요소는 각 소스의 라이선스를 따르며, JSON의 sources[]에 나와 있습니다.
스냅샷은 매월 1일에 확정되며 이후 바뀌지 않습니다. 새 방법론 버전은 새 시리즈를 시작합니다.
월별 JSON: 2026-10
변경 내역
- v1 2026-10-01 First version: LMArena + Epoch AI + τ²-bench, linear equating to Arena text Elo with parameters fixed on the reference snapshot, family shares, sub-scores arena / bench, value = residual of the price regression + Pareto, local = Fit on 5 stacks.
- v1.1 2026-10-01 Local quant floor: the quant is picked like local-fit (Q3 floor; below it only when nothing of Q3+ fits — "strong compression", ×0.6, ranked after every Q3+ model). Local only; the v1 scale and the published 2026-10 snapshot are unchanged.