Leaderboard
Preference ratings from blind human votes, on an Elo-like scale — cost and speed were never visible to voters.
Claude Fable 5
1691
±211
4/7 picked · $1.04/gen
provisional
Kimi K3
1508
±222
2/9 picked · $0.424/gen
provisional
GPT 5.6 Sol
1473
±187
4/15 picked · $0.305/gen
provisional
| # | Model | Rating | W / L | Votes | Avg cost | Median time |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropicprovisional | 1691±211 | 7 | $1.04 | ||
| 2▲1 | Kimi K3Moonshotprovisional | 1508±222 | 9 | $0.424 | ||
| 3 | Claude Opus 5Anthropic | 1500 | 0 | $1.26 | ||
| 4▼2 | GPT 5.6 SolOpenAIprovisional | 1473±187 | 15 | $0.305 | ||
| 5▼1 | GLM 5.2Z.aiprovisional | 1317±236 | 14 | $0.066 |
Ratings come from one preference model fitted to every blind vote at once (Plackett–Luce, on an Elo-like scale where +400 ≈ 10:1 odds). Because it’s fitted globally rather than accumulated vote by vote, a model added later is judged on its own record — not penalised for arriving late. ± is the 95% confidence interval; overlapping intervals mean the gap isn’t established yet. “Provisional” means too few votes to place confidently. W/L is times picked vs times shown. Cost and time are per full generation. ▲/▼ is the 7-day rank move.
Rating vs cost
Up and left is the value frontier — preferred by humans, cheap to run. The leader is highlighted in gold.
Bubble size reflects vote volume.