JOVANA
Explore Library Glossary Getting Started Three Levels Fields How it works Mission
Join the mission
All guides

The Speech BCI Problem: Bandwidth, Anatomy, and Two Paradigms

Why speech restoration is the highest-bandwidth BCI goal, where in cortex it is tractable, and how brain-to-text and direct synthesis divide the field.

Bridging up: from spellers to speech

In Volume I you built the core intuitions: what a brain–computer interface is, how the decoding pipeline turns signals into commands, and what EEG, ECoG and intracortical recording each buy you. This track goes deep on the single most demanding thing a BCI can output: continuous, natural language. The clinical motivation is stark — people living with anarthria from ALS, brainstem stroke, or locked-in syndrome who think in fluent language yet cannot move the muscles that produce speech.

The key conceptual move — the one that makes the problem tractable — is this: a speech neuroprosthesis does not read abstract thoughts. It decodes the motor plan for attempted speech, the same cortical commands that would have driven the vocal tract, still generated even when no sound comes out.

The bandwidth argument

Why pour so much effort into speech rather than a faster speller? Bandwidth. Conversational speech runs at roughly 150 words per minute, and the information content of English text is on the order of one bit per character (Shannon's classic estimate). Put those together and natural speech carries information at well over ten bits per second — an order of magnitude beyond what letter-by-letter P300 or SSVEP spellers deliver in practice.

R_{\text{info}} \;\approx\; \underbrace{2.5}_{\text{words/s}}\;\times\;\underbrace{5}_{\text{char/word}}\;\times\;\underbrace{1}_{\text{bit/char}} \;\approx\; 12\ \text{bits/s}

An order-of-magnitude target: natural speech carries roughly ten-plus bits per second, the bandwidth a speech BCI aspires to.

A back-of-the-envelope estimate: at about 2.5 words a second, 5 characters a word, and roughly one bit of real information per character, natural speech carries on the order of ten-plus bits per second — the target bandwidth a speech BCI chases.

R_{\text{info}}
The information rate, in bits per second.
2.5,\ 5,\ 1
The three factors: words/s, characters/word, and bits/character.

Compare with a good typing-based BCI at a few bits per second — speech aims roughly ten times higher, the goal of a speech neuroprosthesis.

Where decodable speech lives in cortex

Decodable speech lives in the ventral sensorimotor cortex (vSMC), which holds a somatotopic map of the articulators — larynx, lips, tongue, jaw and velum. Crucially this is a motor representation, not a semantic one, which is why it is far more tractable than trying to read meaning out of language-comprehension areas. The distinction between attempted, imagined and inner speech then decides how strong a signal you get: attempted speech, where the participant actively tries to articulate, drives vSMC robustly even with no audible output; inner or covert speech is weaker, more variable, and far less spatially organized.

Two paradigms: text vs voice

Brain-to-text decodes cortex into a discrete symbol stream — characters, phonemes or words. Because the output is language, you can bring a language model into the loop to fix errors and enforce fluency, which makes text the accuracy champion. The price is latency and the loss of everything a transcript throws away: your voice, your prosody, your timing.

Direct speech synthesis instead decodes straight to an audio waveform through a neural vocoder. It can restore a voice with prosody and a real-time, streaming feel — but it is harder and noisier, because you must reconstruct a rich acoustic signal rather than pick from a finite alphabet.

Both paradigms instantiate the same closed loop — acquisition, filtering, feature extraction, decoder — and diverge only at the effector (a text renderer or an audio synthesizer) and in whether feedback is closed in real time.

Which signals can carry speech?

What signal can actually carry speech? High-density ECoG over vSMC is the workhorse: it combines broad coverage with enough spatial and spectral resolution to track the high-gamma band (roughly 70–150 Hz), the local-firing proxy that carries most speech-relevant information. Intracortical microelectrode arrays such as the Utah array add single- and multi-unit resolution but sample a tiny volume. Scalp EEG, by contrast, is too blurred by volume conduction and lacks accessible high-gamma, so continuous speech decoding from EEG remains out of reach.