错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enabling the Translation of Electromyographic Signals Into Speech: A Neural Network Based Decoding Approach

  • Abhishek Bharali,
  • Bidyut Bikash Borah,
  • Uddipan Hazarika,
  • Soumik Roy

摘要

Speech, the principal mode of human interaction, involves the articulation of language through vocal sounds generated by the vocal apparatus. It encompasses various forms such as vocalized speech, whispering, silent speech, and subvocal speech. Silent speech refers to the absence of audible sound despite movement of speech articulators due to minimized airflow. This study seeks to convert silently mouthed words into audible speech aimed at providing communication assistance to individuals with speech impairments. The research utilizes electromyographic (EMG) signals captured from facial muscles during speech production, in conjunction with corresponding audio recordings of the same. Both the EMG signals and the audio recordings are utilized for feature extraction are the extracted features are then collectively employed to train three distinct neural network models viz. Convolutional neural network (CNN) model, Gated Recurrent Unit (GRU) and Convolutional neural network Long Short Term Memory (CNN-LSTM). Model predicts the audio features based on EMG features input and they are subsequently passed through vocoder to reconstruct the original audio speech. The models are tested on real time data and the corresponding metrics and plots are evaluated. The performance metrics establishes the superiority of the CNN-LSTM model over the other models with mean squared error (MSE) as low as 0.036. Such an approach holds promise for improving communication aids and speech rehabilitation technologies.