Scaling Laws & Emergence

the emergence debate

If a skill seems to appear from nowhere at a certain size, a natural question follows: is the jump really in the model, or in our ruler? The emergence debate is the live argument over exactly this. One camp sees genuine, phase-change-like leaps; another argues many emergent curves are mirages created by harsh, all-or-nothing scoring.

The skeptical case: if you grade a task as fully right or fully wrong — exact-match accuracy on a long answer — then steady, smooth gains in the model's underlying probability of the correct sequence stay invisible until they cross a threshold, at which point the score suddenly leaps. Swap to a smoother metric, like token-level likelihood or partial credit, and the same runs often reveal a gradual, predictable climb with no magic step. The counter-argument: some abilities really do look discontinuous even under gentle metrics, and predictable in hindsight is not the same as predictable in advance.

The stakes are practical, not philosophical: if sharp emergence is mostly a metric artefact, capabilities are more forecastable and safer to plan around; if it is real, surprises are baked into scaling.