Evaluation & Benchmarks

MMLU-style benchmarks

These are big multiple-choice exams that sweep across human knowledge — history, law, medicine, physics, ethics, and dozens more subjects — to test how much a model knows and can recall under pressure. MMLU (Massive Multitask Language Understanding) is the famous example: thousands of four-option questions drawn from real coursework and professional tests. The model reads a question, picks A, B, C, or D, and we score it simply by accuracy across all subjects.

Their appeal is breadth and easy grading. One number summarises performance over 57 fields, and because answers are single letters there is no ambiguity in scoring. They became the default headline metric for general capability, and successors like MMLU-Pro add harder questions and more options to fight saturation.

Their weakness is that multiple choice rewards recognising the right answer, not producing it, and a lucky guess scores the same as real understanding. They also leak readily into training data and can contain mislabelled items, so a top MMLU score signals broad knowledge but says little about reasoning, honesty, or whether the model can do anything useful unprompted.

Also called
multiple-choice knowledge examsMMLU