live benchmark report
Frontier AI models, judged without the badge
Every model receives the same creative brief. People inspect the resulting websites, games, visualizations and tools without seeing the model names, then pick one winner. This report turns those choices into a live preference ranking and keeps cost and speed beside quality.
Blind votes
40
Comparisons
152
Models
5
Challenges
15
What the evidence says now
Claude Opus 5 currently leads the overall preference ranking at 1704points, based on 1 blind appearances. The result remains provisional while more votes accumulate.
GLM 5.2 has the lowest observed generation cost in the rated field at roughly $0.066 per output.
GPT 5.6 Sol is the fastest rated model by median generation time at 1m 32s.
Leaders by task
| Category | Current leader | Rating | Evidence |
|---|---|---|---|
| Web design | Kimi K3 | 1580 | 27 appearances |
| Game | Claude Fable 5 | 1574 | 9 appearances |
| Animation | Claude Fable 5 | 1691 | 7 appearances |
| Data viz | GLM 5.2 | 1650 | 3 appearances |
| Tool | Claude Fable 5 | 1653 | 6 appearances |
| Generative | Claude Opus 5 | 1694 | 1 appearances |
How to read this report
Ratings describe human preference, not factual correctness or hidden code quality. Confidence intervals and provisional labels matter: small samples can move quickly. Costs are the billed API cost for the stored generation, and voters never see price, latency or identity before choosing.
Explore the full leaderboardRead the methodologyContribute a blind vote