Models
Every model on the bench gets the same briefs and is judged blind. Each profile shows its human-preference rating, its record, and the receipts — what its outputs cost and how long they took.
Claude Opus 5
Anthropicprovisional1623±313
picked 1/2 · $1.26/gen · 8m 55s
Full profileClaude Fable 5
Anthropic1557±96
picked 20/60 · $1.04/gen · 3m 22s
Full profileGPT 5.6 Sol
OpenAI1486±97
picked 18/72 · $0.305/gen · 1m 32s
Full profileKimi K3
Moonshot1481±101
picked 15/66 · $0.424/gen · 10m 38s
Full profileGLM 5.2
Z.ai1432±104
picked 13/73 · $0.066/gen · 3m 11s
Full profileRatings are blind human preference (Plackett–Luce, Elo-like scale); ± is the 95% confidence interval. See the methodology for how these are computed.