Skip to content

Leaderboard

Preference ratings from blind human votes, on an Elo-like scale — cost and speed were never visible to voters.

1Anthropic

Claude Opus 5

1704

±360

1/1 picked · $1.26/gen

provisional

2Anthropic

Claude Fable 5

1533

±104

19/59 picked · $1.04/gen

3OpenAI

GPT 5.6 Sol

1472

±103

18/71 picked · $0.305/gen

#ModelRatingVotesAvg cost
1▲4
Claude Opus 5Anthropicprovisional
1704±360
1$1.26
2▼1
Claude Fable 5Anthropic
1533±104
59$1.04
3▼1
GPT 5.6 SolOpenAI
1472±103
71$0.305
4▼1
Kimi K3Moonshot
1467±107
65$0.424
5▼1
GLM 5.2Z.ai
1417±110
72$0.066

Ratings come from one preference model fitted to every blind vote at once (Plackett–Luce, on an Elo-like scale where +400 ≈ 10:1 odds). Because it’s fitted globally rather than accumulated vote by vote, a model added later is judged on its own record — not penalised for arriving late. ± is the 95% confidence interval; overlapping intervals mean the gap isn’t established yet. “Provisional” means too few votes to place confidently. W/L is times picked vs times shown. Cost and time are per full generation. ▲/▼ is the 7-day rank move.

Rating vs cost

Up and left is the value frontier — preferred by humans, cheap to run. The leader is highlighted in gold.

14251500157516501725$0.05$0.1$0.2$0.5$1.00$2.00Avg cost per generation (log scale)RatingClaude Opus 5Claude Fable 5GPT 5.6 SolKimi K3GLM 5.2

Bubble size reflects vote volume.