Skip to content

live benchmark report

Frontier AI models, judged without the badge

Every model receives the same creative brief. People inspect the resulting websites, games, visualizations and tools without seeing the model names, then pick one winner. This report turns those choices into a live preference ranking and keeps cost and speed beside quality.

Blind votes

40

Comparisons

152

Models

5

Challenges

15

What the evidence says now

Claude Opus 5 currently leads the overall preference ranking at 1704points, based on 1 blind appearances. The result remains provisional while more votes accumulate.

GLM 5.2 has the lowest observed generation cost in the rated field at roughly $0.066 per output.

GPT 5.6 Sol is the fastest rated model by median generation time at 1m 32s.

Leaders by task

CategoryCurrent leaderRatingEvidence
Web designKimi K3158027 appearances
GameClaude Fable 515749 appearances
AnimationClaude Fable 516917 appearances
Data vizGLM 5.216503 appearances
ToolClaude Fable 516536 appearances
GenerativeClaude Opus 516941 appearances

How to read this report

Ratings describe human preference, not factual correctness or hidden code quality. Confidence intervals and provisional labels matter: small samples can move quickly. Costs are the billed API cost for the stored generation, and voters never see price, latency or identity before choosing.

Explore the full leaderboardRead the methodologyContribute a blind vote