Cross-modal alignment and fusion of EEG-visual based on mixed attention mechanism for emotion recognition
摘要
Changes in human emotions are often accompanied by changes in multiple physiological or external stimuli. Fusing these multi-modal information can improve the accuracy of emotion recognition. However, since current multi-modal emotion recognition algorithms do not consider modal synchronization, the resulting misalignment of information affects emotion recognition accuracy. To address these issues, this paper proposes an Electroencephalogram (EEG)-visual cross-modal alignment and fusion model for emotion classification (CMAF) based on a hybrid attention mechanism. Differential entropy (DE) features of five frequency bands of EEG signals and 10 color features of visual modalities are utilized for cross-modal emotion recognition. The cross-modal alignment module extracts key information by a multi-head attention mechanism, and improves the similarity of two modes under the constraint of loss function. A cross-attention module is designed to use visual modalities to guide feature extraction of EEG signals and establish correlations between two modalities for modal fusion. Support vector machine (SVM) is used to classify the features extracted from different emotional states in SEED dataset. Experimental results show that fusing high-frequency EEG and video features significantly improves recognition accuracy, with the Gamma-visual fusion achieving an average accuracy of 96.49%. To further evaluate the model’s generalization capability, we introduced the SEED-IV dataset and conducted assessments on two datasets under both subject-related and subject-independent settings. The results demonstrate that the model consistently maintains robust performance across diverse data sources, highlighting its robustness and generalization potential.