Skip to content

Models

Every model on the bench gets the same briefs and is judged blind. Each profile shows its human-preference rating, its record, and the receipts — what its outputs cost and how long they took.

Ratings are blind human preference (Plackett–Luce, Elo-like scale); ± is the 95% confidence interval. See the methodology for how these are computed.