Foundation Models & Representation Learning for Neural Data

Scaling laws for neural data

The question of whether representation or decoding quality improves predictably — ideally as a smooth power law — as you increase the amount of neural data (neurons, hours, subjects), model size, and compute, mirroring the scaling laws found for language and vision (Kaplan-style and Hoffmann/Chinchilla-style curves). If such laws held, they would license the whole foundation-model strategy: buy performance with scale and know in advance what a bigger dataset buys.

The honest status is that this is a motivating hypothesis with only partial, preliminary empirical support. Neural datasets are orders of magnitude smaller and far more heterogeneous than internet-scale corpora, and adding subjects can inject distribution shift rather than clean, in-distribution scale. Early studies report that more sessions and more neurons help, but whether the improvement follows a clean, extrapolatable power law — and where it saturates — is unresolved. Treat published scaling curves as suggestive, not as established law.

Also called
neural scaling lawsdata-scaling behavior