benchmark saturation
A benchmark saturates when the best models score so high — near the ceiling — that it can no longer tell them apart. If three models all hit 96, 97, and 98 percent, the differences are within noise and mislabelled items, and the benchmark has stopped doing its job: ranking the frontier. The remaining headroom is too small to be meaningful.
Saturation is partly success — models really did get better — and partly an artefact, since the easy questions get solved, contamination inflates scores, and any errors left in the test set cap the achievable maximum. Either way, a saturated benchmark gives a false sense that progress has stalled or that two very different models are equivalent.
The field responds by retiring saturated sets and building harder successors — more difficult questions, more answer options, adversarially filtered items, expert-level or contamination-resistant tasks. The treadmill is permanent: today's impossible benchmark is next year's saturated one, so evaluation must keep moving the goalposts to stay informative.