Models
Every model on the bench gets the same briefs and is judged blind. Each profile shows its human-preference rating, its record, and the receipts — what its outputs cost and how long they took.
Claude Opus 5
Anthropicprovisional1704±360
picked 1/1 · $1.26/gen · 8m 55s
Full profileClaude Fable 5
Anthropic1533±104
picked 19/59 · $1.04/gen · 3m 22s
Full profileGPT 5.6 Sol
OpenAI1472±103
picked 18/71 · $0.305/gen · 1m 32s
Full profileKimi K3
Moonshot1467±107
picked 15/65 · $0.424/gen · 10m 38s
Full profileGLM 5.2
Z.ai1417±110
picked 13/72 · $0.066/gen · 3m 11s
Full profileRatings are blind human preference (Plackett–Luce, Elo-like scale); ± is the 95% confidence interval. See the methodology for how these are computed.