Agents & Frontier

scaling laws (capability)

/ SKAY-ling lawz /

Scaling laws are a surprising empirical pattern: when you train ever-larger AI models on ever-more data with ever-more computing power, their performance improves in a smooth, predictable way — not in random leaps, but along a steady curve you can plot and extrapolate. It's a bit like discovering that, on average, a bigger engine reliably produces proportionally more power: the relationship is regular enough to plan around.

More precisely, researchers found around 2020 that a model's prediction error tends to fall in a tidy mathematical relationship (a power law) as you increase three ingredients — model size, dataset size, and compute. This was a big deal because it let labs forecast roughly how good a model would be before spending millions to train it, and it largely explains the strategy of the past few years: pour in more scale and capability climbs along the curve. The word "law" here means a robust observed trend, not a law of nature like gravity.

Two honest caveats matter. First, the smooth curve usually measures a low-level statistic (how well the model predicts the next token), which is not the same as the practical abilities we care about — those can appear more jaggedly. Second, scaling is not free or infinite: it costs enormous money, energy, and data, and there are real worries that high-quality training data is running short and that returns are diminishing. Scaling laws describe what has happened so far; they are not a guarantee that piling on more will keep paying off, or that scale alone is the road to general intelligence.

A lab plans a new model. Using scaling-law curves fit to smaller training runs, it estimates that a model 10× larger trained on 10× more data should cut prediction error by a predictable amount — so it commits the budget before training, expecting to land near the forecast. This forecasting power, not magic, is why scale became the dominant strategy.

Scaling laws let labs forecast a model's loss before paying to train it.

"Law" overstates it. Scaling laws are observed regularities over a particular range, not guarantees that hold forever — they can bend or break as data quality, model types, or what you measure change. A trend that has held so far is evidence, not a promise about the next order of magnitude.

Also called
scaling lawsneural scaling laws缩放定律縮放定律尺度定律