ProCFD: Towards Robust Multimodal Sentiment Analysis Through Prototype Fusion and Contrastive Feature Decomposition
摘要
When faced with complex emotional data, existing models struggle to achieve complete feature decomposition, resulting in insufficient discrimination of the generated decomposition features. This is primarily due to the unnecessary cross-feature interference. In addition, the diversity and ambiguity of emotion expression in complex data further aggravate the interference between features. In this paper, we put forward a multimodal sentiment analysis framework, ProCFD, which comprises three key components: the prototype fusion module, the feature decomposition module, and the contrastive learning mechanism. The designed prototype fusion module generates more discriminative enhanced text features through prototype enhancement and employing two different audio-visual fusion strategies in line with the datasets, significantly improving the representational capability of multimodal features. The proposed feature decomposition module decomposes multimodal features into similar and dissimilar features and uses an autoregressive approach to regenerate dissimilar features, effectively reducing cross-feature interference. Our contrastive learning mechanism is achieved through designing innovative inter-sample and intra-sample contrast, which effectively strengthen the model’s resilience to complex emotional information. Experiments on the commonly used datasets SIMS, MOSI, and MOSEI demonstrate that ProCFD outperforms existing methods across all benchmark metrics, validating its effectiveness in processing complex emotional data and enabling robust sentiment analysis.