Articulatory (kinematic) speech decoding
An intermediate-representation strategy: instead of mapping neural activity straight to sound or text, first decode the movements of the vocal-tract articulators, lip aperture, tongue body and tip position, jaw opening, and larynx and voicing, often expressed as a gestural score or articulatory-feature trajectory, and only then map those movements to acoustics. The rationale is that motor cortex encodes movement, so articulatory kinematics are more linearly related to neural activity than the raw acoustic spectrogram is.
Anumanchipalli, Chartier and Chang (2019) used exactly this two-stage neural-to-articulatory-to-acoustic pipeline to synthesize intelligible sentences from ECoG, and showed that the decoded articulatory representation generalized better than a direct neural-to-acoustic mapping.
Decode a trajectory in which lip aperture closes and then reopens while the larynx is voiced, the articulatory signature of a /b/, and hand it to a synthesizer, rather than trying to classify the acoustic burst directly.
Articulatory intermediates make motor-cortex signals easier to map than acoustics.
In a paralyzed user the true articulator movements during attempted speech cannot be observed, so they are inferred with an acoustic-to-articulatory inversion model fitted on able-bodied speakers; this borrowed kinematic target introduces model dependence and possible mismatch.