Cross-Modal Sentiment Analysis Based on Fine-Grained Feature Interaction Learning
摘要
Driven by the advancement of mobile internet technology, people are no longer confined to using only text to convey information; instead, they increasingly prefer a combination of different modalities. Among these, the integration of text and images is the most common. Although cross-modal sentiment analysis has more sources of information compared to unimodal sentiment analysis, it faces new issues and challenges. To address the issue of existing methods failing to accurately extract and fuse fine-grained image and text features, we propose a cross-modal sentiment analysis method based on fine-grained features (CSAFGF). By deeply exploring the interactive information between image and text features, the method can more comprehensively capture and integrate fine-grained emotional features across different modalities, thereby enhancing the model’s performance. To validate the effectiveness of the proposed method, we conducted experiments on various datasets for cross-modal sentiment analysis involving text and images. The experimental results demonstrate that the proposed CSAFGF method shows improvements in evaluation metrics compared to other models.