EGSRNet: Emotion-Label Guiding and Similarity Reasoning Network for Multimodal Sentiment Analysis
摘要
Multimodal sentiment analysis has attracted many research interests in social media. Existing methods mainly rely on mining the global/local information in the image, to realize the fusion with better text information, while ignoring the inherent semantic information contained in the text. For this purpose, the Emotion-label Guiding and Similarity Reasoning Network(EGSRNet) is proposed, which introduces emotion-label guided text features to extract hidden semantic information to improve local image-text interaction, and realize deeper understanding and analysis of image-text by combining context information. Specifically, the Image-Text Feature Extraction module is used to fully extract the global/local-entity image-text features to improve the utilization rate of vital features. For text features, the emotion-label is introduced to enhance the representation ability of deep semantic information. Secondly, to explicitly calculate the similarity between text and local-entity image features, capture the image-text correlation and fully interact, a Local-Entity Similarity Reasoning module based on the attention mechanism is designed. Finally, multimodal interaction is achieved by combining the global image-text context, and the data/label-based contrastive learning is introduced to improve performance. Experimental results show that the proposed model outperforms the baseline methods on three public datasets.