A method for multimodal sentiment analysis: adaptive interaction and multi-scale fusion
摘要
To address the potential issue of introducing irrelevant emotional data during the fusion of multimodal representations and the challenge of neglecting critical information at different scales within the sequence information after fusion, this study innovatively proposes the Adaptive Interaction and Multi-Scale Fusion Model (AIMS). Initially, multimodal feature fusion is achieved through the interaction of text features with video features and audio features respectively. Subsequently, two feature vectors related to text are subjected to deep interactive fusion to generate a comprehensive modal representation, thereby enhancing the expression of text-related information. Then, the model clearly employs a multi-scale feature pyramid network structure for feature extraction at multiple scales, effectively fusing sequence information in the integrated modal representation and capturing key sequence features in multimodal data. Finally, the multimodal fusion module integrates the final modal representation for more effective Multimodal Sentiment Analysis(MSA). Experimental results demonstrate that the AIMS model outperforms existing sentiment analysis models on the CMU-MOSEI, CMU-MOSI, CH-SIMSv2 and MELD datasets.