Integrating Time-Frequency Domain Shallow and Deep Features for Speech-EEG Match-Mismatch of Auditory Attention Decoding
摘要
Electroencephalogram (EEG) signals provide an important pathway to reflect brain activations, from which auditory attention clues of the listener can be decoded, termed as auditory attention decoding (AAD). However, existing AAD methods primarily rely on temporal or frequency features of audio and shallow features of EEG. In this work, we propose a new model fusion based AAD method with residual dilated convolution blocks, which considers both shallow and deep attention mechanisms as well as time-frequency domain features. Besides, EEG data from different utterances are mixed with the selected EEG segment for augmentation to increase the sample diversity. The effectiveness of our approach is verified by the match-mismatch task of ICASSP2024 Auditory EEG Challenge, which is a typical example of AAD. It performs much better than the baseline and state-of-the-art methods.