Foundation Models & Representation Learning for Neural Data

Large-scale neural data corpus

The pretraining-data problem: assembling neural recordings large and diverse enough to train a foundation model. Public efforts point the way — the International Brain Laboratory standardized Neuropixels recordings across many labs, human intracranial collections (AJILE and similar), and large clinical EEG corpora such as the Temple University Hospital archive — and standardized formats like NWB (Neurodata Without Borders) make pooling feasible.

Even so, the largest neural corpora are minuscule beside the text and image datasets that trained the models the field takes as inspiration. Heterogeneity cuts both ways: variety across species, areas, tasks and rigs is what a general model needs, yet it is also distribution shift that can blunt or reverse transfer. Data sharing, consent, standardization and the sheer cost of chronic recording are the practical rate-limiters, which is why data aggregation, not model architecture, is often the true bottleneck to scaling.

Also called
neural pretraining datasetdata aggregation for BCI