Identification of Speech Stream and the Source Localization for Hearing Prosthesis-Driven Healthcare
摘要
The human auditory system demonstrates the capacity to selectively attend to a single speaker amidst multiple speakers and environmental noise within a speech stream. In such instances, cortical activity exhibits a heightened synchronization with the envelope of the uttered speech compared to the unspoken discourse. Utilizing electroencephalography (EEG) signals enables the inference of the attended speaker by discerning neural activity patterns in relation to the speech signals. Given the nonlinear nature of the human brain, employing deep learning methodologies becomes instrumental in disentangling the dynamic state of the brain from neural signals. However, many nonlinear methods proposed in recent research either neglect the incorporation of speech streams or rely solely on the speech envelope as input features. Enhancing the performance of deep learning can be achieved by including information pertaining to the streams of speech. We introduced a unified model consisting of a Convolutional Neural Network (CNN) and Bilinear-Long Short-Term Memory (BiLSTM) architecture to extract auditory attention. This model is designed to leverage cortical recordings and spectrograms from multiple speech streams for decoding attention in scenarios involving two speakers.