Mtfsfn: a multi-view time-frequency-space fusion network for EEG-based emotion recognition
摘要
Over the recent years, emotion recognition based on electroencephalogram (EEG) has emerged as a prominent research area. Nevertheless, EEG signals present spatially discrete and non-stationary characteristics, to represent spatiotemporal information and extract more discriminative features from complex signals is still a challenge. This study proposed a multi-view time-frequency-space fusion network, referred to as MTFSFN. To effectively utilize complementary information from different frequency bands, we employ a frequency-domain attention mechanism to allocate weights to features of different frequency bands. A multi-view Transformer model was designed, integrating Transformer with two-dimensional positional embeddings to extract discrete spatial information. Following the fusion of multi-view features, we utilize LSTM to capture dynamic time-frequency-space relationships. Finally, a subject-independent leave-one-subject-out cross-validation strategy was used to validate extensively on three public datasets, DEAP, SEED, and SEED-IV. On the DEAP dataset, the average accuracies of valence and arousal are 78.64% and 77.42%, respectively. On the SEED dataset, the average accuracy is 86.91%. On the SEED-IV dataset, the average accuracy is 75.51%. The experimental results show that the proposed MTFSFN model achieves excellent recognition performance.