Foundation Models & Representation Learning for Neural Data

Few-shot / cross-session decoding benchmark

Standardized evaluations built to test the exact promise of a foundation model — transfer with little calibration — rather than convenient in-distribution accuracy. FALCON (Few-shot Algorithms for Consistent Neural decoding) measures how well a decoder holds up on held-out days and subjects with only a small calibration set, across several implant modalities. The Neural Latents Benchmark (NLB) evaluates the quality of inferred latent states across a suite of datasets with common metrics such as co-smoothing. These benchmarks make claims comparable and reproducible.

Their value is disciplinary. A model earns the label 'foundation' only if it transfers on data it never saw — new sessions, new subjects — with minimal target data, which is precisely what these benchmarks isolate. Without such held-out, few-shot evaluation, strong-looking numbers can reflect memorized in-distribution structure rather than the generality the field is trying to buy.

In-distribution decoding accuracy and held-out cross-session transfer are different quantities; report both, and never present the former as evidence for the latter.

Also called
FALCON benchmarkNeural Latents Benchmarktransfer benchmark for BCI