본문으로 건너뛰기
fedi.software

fedi-index 방법론

지수의 소스, 라이선스, 가중치와 공식 — 실제 계산에 쓰이는 것과 같은 설정에서 생성됩니다.

방법론 v1 · 2026-10 실시간 데이터(매일 재계산) · 추이 · JSON

  • 명시적인 오픈 라이선스(CC BY, MIT, Apache)가 있는 소스만 지수에 반영합니다. 허가를 기다리는 소스는 꺼 둡니다.
  • 독립적인 실행 결과만 반영합니다. 제조사 자체 발표는 표시와 함께 보여 주되 가중치를 주지 않습니다.
  • 소스 계열 하나는 한 표: 같은 투표자의 상관된 카테고리는 합산하지 않습니다.
  • 모든 점수의 구성이 공개됩니다: 구성 요소, 소스, 값, 가중치, 날짜.
  • 불확실성을 드러냅니다: 점수에는 구간이 붙고, 데이터가 적으면 ‘잠정’으로 표시합니다.

소스와 라이선스

소스 계열 라이선스 지수 반영
LMArena Arena Leaderboard Dataset by LMArena, CC BY 4.0. Modified by fedi.software (equated, weighted, aggregated). Not endorsed by LMArena. Arena(사람 투표) CC BY 4.0 예
Epoch AI Epoch AI, 'Capabilities & benchmarking'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks'. CC BY. ECI by Epoch AI; Epoch-run results only. Modified by fedi.software. Epoch AI(벤치마크) CC BY 4.0 예
τ²-bench (Sierra) τ²-bench leaderboard, © Sierra Research, MIT License (Yao et al. 2024, arXiv:2406.12045; Barres et al. 2025, arXiv:2506.07982). Sierra-run text submissions only. τ²-bench(에이전트) MIT 예
LiveBench LiveBench — 허가 대기 중
Terminal-Bench Terminal-Bench — 허가 대기 중
SWE-rebench SWE-rebench — 허가 대기 중
Toolathlon Toolathlon — 허가 대기 중
models.dev Pricing: models.dev, MIT License, © 2025 models.dev. modelsdev MIT 예
LiteLLM Pricing: LiteLLM model_prices_and_context_window.json, MIT License, © 2023 Berri AI. litellm MIT 예
Hugging Face Downloads (30 days): Hugging Face Hub, as of the snapshot date. hf — 예

분야와 가중치

종합

Arena(사람 투표) 비중 50%

  • text/overall text/overall2
  • text/hard_prompts text/hard_prompts2
  • text/expert text/expert1
  • text/instruction_following text/instruction_following1
  • text/multi_turn text/multi_turn0.50
  • text/longer_query text/longer_query0.50

Epoch AI(벤치마크) 비중 50%

  • Epoch Capabilities Index eci3
  • GPQA Diamond gpqa_diamond0.50
  • SimpleQA Verified simpleqa_verified0.50

코딩

Arena(사람 투표) 비중 60%

  • text/coding text/coding1.5
  • webdev/overall webdev/overall1.5

Epoch AI(벤치마크) 비중 40%

  • SWE-bench Verified swe_bench_verified1

에이전트

Arena(사람 투표) 비중 60%

  • agent/overall agent/overall1

τ²-bench(에이전트) 비중 40%

  • τ-Knowledge banking tau2/banking_knowledge1.5
  • τ²-bench retail tau2/retail1
  • τ²-bench airline tau2/airline1
  • τ²-bench telecom tau2/telecom1

수학

Arena(사람 투표) 비중 50%

  • text/math text/math1

Epoch AI(벤치마크) 비중 50%

  • FrontierMath Tiers 1–3 frontiermath_t1_32
  • FrontierMath Tier 4 frontiermath_t41

비전 잠정

Arena(사람 투표) 비중 100%

  • vision/overall vision/overall2
  • vision/ocr vision/ocr0.50
  • vision/diagram vision/diagram0.50

검색 잠정

Arena(사람 투표) 비중 100%

  • search/overall search/overall1

글쓰기 잠정

Arena(사람 투표) 비중 100%

  • text/creative_writing text/creative_writing2
  • text/longer_query text/longer_query1

언어 잠정

언어별 Arena 투표; 해당 분야는 그 언어 버전 사이트에 표시됩니다.

English — text/english · Русский — text/russian · Deutsch — text/german · Español — text/spanish · Français — text/french · 日本語 — text/japanese · 한국어 — text/korean · 繁體中文 — text/chinese · Polski — text/polish

벤치마크와 경과 기간

구성 요소 버전 출시
Epoch Capabilities Index eci 실시간(감쇠 없음)
GPQA Diamond gpqa_diamond 2023-11-20
SimpleQA Verified simpleqa_verified 2025-09-09
FrontierMath Tiers 1–3 frontiermath_t1_3 2026-06-12
FrontierMath Tier 4 frontiermath_t4 2026-06-12
SWE-bench Verified swe_bench_verified 2024-08-13
τ²-bench retail tau2/retail 2024-06-17
τ²-bench airline tau2/airline 2024-06-17
τ²-bench telecom tau2/telecom 2025-06-09
τ-Knowledge banking tau2/banking_knowledge 2026-03-04

오래된 벤치마크 버전일수록 가중치가 낮아집니다(학습 데이터 오염, 포화). 가중치는 decay_half_life개월마다 절반이 되며 decay_min 아래로는 내려가지 않습니다.

공식

각 구성 요소는 양쪽에서 모두 측정된 모델을 기준으로 선형 동등화를 거쳐 Arena 척도로 옮깁니다(앵커 A = Arena text overall). 매개변수는 방법론 버전의 기준 스냅샷에 고정되므로 월별 비교가 가능합니다.

x_eq  = μ_A + (x − μ_c) · σ_A / σ_c          se_eq = se · σ_A / σ_c
w_c   = 8 · family_share[f] · base_w[c] / Σ base_w(f, present)
rel_c = τ² / (τ² + se_eq²)
w_eff = w_c · rel_c · decay_c · (self_report ? 0 : 1)
Q     = (Σ w_eff · x_eq + w0 · prior) / (Σ w_eff + w0)
SE(Q) = √Σ (w_eff · se_eq)² / (Σ w_eff + w0)

분야 안에서 구성 요소는 가중치, 신뢰도(τ), 벤치마크의 경과 기간에 따라 평균합니다. 데이터가 적은 모델은 사전값 쪽으로 당겨집니다. 매개변수:

anchor
arena:text/overall
total_weight
8
tau
15
w0
0.50
prior_pct
25
min_anchors
8
min_votes
300
min_confidence
0.90
decay_half_life
18
decay_min
0.25
outlier_sigma
3
outlier_factor
0.50
population_months
18
jump_sigma
3

등급 = 백분위 기준 등급과 신뢰 구간 기준 등급 중 낮은 쪽; 소스 계열이 하나뿐인 모델은 최고 A입니다.

등급 = 백분위 기준 등급과 신뢰 구간 기준 등급 중 낮은 쪽; 소스 계열이 하나뿐인 모델은 최고 A입니다.

provisional_max
A

S ≥ p95A ≥ p80B ≥ p50C ≥ p25 D < p25

신뢰 구간 기준 등급(순위 상한 / 모델 수): S ≤ 10% · A ≤ 30% · B ≤ 60% · C ≤ 85%

가성비: 가격 대비 품질

가성비는 가격의 log₂에 대한 품질 회귀선보다 얼마나 위에 있는지입니다. 가격은 제공업체 중앙값입니다.

slice
general
input_share
3
output_share
1
price_sources
modelsdev, litellm
unit
1m_tokens

내 하드웨어에서

로컬: 오픈 웨이트만, 각 환경에 적재되는 최상의 양자화를 쓰고 점수를 생성 속도로 보정합니다.

slice
general
context
8,192
speed_ref
20
low_penalty
0.60
  • GPU 12 GB
  • GPU 24 GB
  • Mac 64 GB
  • 통합 메모리 128 GB
  • CPU, RAM 32 GB

데이터, 라이선스, 아카이브

지수(자체 수치와 구성 내역)는 CC BY 4.0입니다: “fedi-index by fedi.software”와 링크를 표기하세요. 구성 요소는 각 소스의 라이선스를 따르며, JSON의 sources[]에 나와 있습니다.

스냅샷은 매월 1일에 확정되며 이후 바뀌지 않습니다. 새 방법론 버전은 새 시리즈를 시작합니다.

월별 JSON: 2026-10

변경 내역

  • v1 2026-10-01 First version: LMArena + Epoch AI + τ²-bench, linear equating to Arena text Elo with parameters fixed on the reference snapshot, family shares, sub-scores arena / bench, value = residual of the price regression + Pareto, local = Fit on 5 stacks.
  • v1.1 2026-10-01 Local quant floor: the quant is picked like local-fit (Q3 floor; below it only when nothing of Q3+ fits — "strong compression", ×0.6, ranked after every Q3+ model). Local only; the v1 scale and the published 2026-10 snapshot are unchanged.