Scaling Laws & Emergence

broken scaling laws

A scaling law is a comforting straight line — until it isn't. Sometimes the loss-versus-scale curve does not stay a single tidy slope: it bends, kinks, or even briefly gets worse before improving again. Broken scaling laws are the name for these departures from one clean power law, and they matter because a naive extrapolation across a break can badly mislead a billion-dollar training decision.

The richer description is a smoothly-broken power law: several power-law segments with different slopes, stitched together by gentle transitions, so the curve can steepen, flatten, or change regime as scale grows. Breaks show up where a new capability switches on, where a model exhausts one kind of structure and starts exploiting another, or where the data or optimisation hits a wall. Some are predictable once you fit the broken form; others, like a curve that plateaus then drops, are genuinely hard to foresee from the smaller side of the bend.

Breaks cut both ways — a flattening can wrongly suggest scaling is dead just before a new regime ignites, and a steep early slope can flatter a forecast that later stalls. They are the main reason extrapolation carries risk.

Also called
smoothly broken power laws