Skip to content

Leaderboard

Preference ratings from blind human votes, on an Elo-like scale — cost and speed were never visible to voters.

1Z.ai

GLM 5.2

1650

±290

2/3 picked · $0.066/gen

provisional

2OpenAI

GPT 5.6 Sol

1524

±303

1/3 picked · $0.305/gen

provisional

3Moonshot

Kimi K3

1490

±290

1/4 picked · $0.424/gen

provisional

#ModelRatingVotesAvg cost
1▲3
GLM 5.2Z.aiprovisional
1650±290
3$0.066
2
GPT 5.6 SolOpenAIprovisional
1524±303
3$0.305
3
Claude Fable 5Anthropic
1500
0$1.04
4
Claude Opus 5Anthropic
1500
0$1.26
5▼2
Kimi K3Moonshotprovisional
1490±290
4$0.424

Ratings come from one preference model fitted to every blind vote at once (Plackett–Luce, on an Elo-like scale where +400 ≈ 10:1 odds). Because it’s fitted globally rather than accumulated vote by vote, a model added later is judged on its own record — not penalised for arriving late. ± is the 95% confidence interval; overlapping intervals mean the gap isn’t established yet. “Provisional” means too few votes to place confidently. W/L is times picked vs times shown. Cost and time are per full generation. ▲/▼ is the 7-day rank move.

Rating vs cost

Up and left is the value frontier — preferred by humans, cheap to run. The leader is highlighted in gold.

1500155016001650$0.05$0.1$0.2$0.5Avg cost per generation (log scale)RatingGLM 5.2GPT 5.6 SolKimi K3

Bubble size reflects vote volume.