Elo leaderboards
Elo leaderboards turn thousands of head-to-head battles into a single rating per model, borrowing the system invented to rank chess players. Each model has a number; when one wins a match-up its rating rises and the loser's falls, with bigger moves for upsets against a stronger opponent. After enough games the ratings settle into a ranking where the gap between two numbers predicts how often one model beats the other.
The most famous example collects real anonymous votes from users who chat with two unnamed models and pick the better reply, then aggregates them into a live leaderboard. Because the prompts come from real people and the judges are real users, this captures broad, everyday preference in a way fixed benchmarks cannot, and it resists memorisation since the questions are never published in advance.
The caveats: Elo measures preference, not truth, so a charming, confident, well-formatted answer can outrank a more accurate dull one. Votes are noisy, popular models get more battles and tighter ratings, and the crowd's taste is not your task. Read it as a popularity-weighted relative ranking, not an absolute measure of capability or safety.
Elo's expected win probability of model A over B from their rating gap; a 400-point lead predicts a 10-to-1 edge.