A High-Performance Neuroprosthesis for Speech Decoding and Avatar Control
摘要
Speech neuroprostheses hold the promise of reinstating communication abilities in individuals with paralysis, yet achieving naturalistic speed and expressiveness remains challenging (Moses, The New England Journal of Medicine 385:217, 2020). In our study (Metzger, Nature 620:1037, 2023), we utilized high-density surface recordings from the speech cortex of a participant in a clinical trial who had significant paralysis in limbs and voice. This enabled high-performance real-time decoding in three distinct speech-related outputs: text, speech audio, and facial-avatar animation. We trained deep-learning models on neural data gathered while the participant tried to silently articulate sentences. For text, our results showed high-speed, large-vocabulary decoding with a median of 78 words per minute and a 25% median word error rate. In terms of speech audio, we achieved intelligible speech synthesis that could be personalized to the participant’s pre-injury voice. For facial-avatar animation, we successfully decoded virtual movements of the mouth and face for both speech and non-verbal communication gestures. The decoders attained high efficiency with less than 2 weeks of training. Our research presents a novel multimodal speech-neuroprosthetic method that offers significant potential to restore comprehensive, embodied communication for individuals with severe paralysis.