A Text-Oriented Transformer with an Image Aesthetics Assessment Fusion Network for Visual-Textual Sentiment Analysis
摘要
The rapid advancement of social networks has significantly altered how people convey their emotions, increasingly through a mix of images and text on social media platforms. Visual-textual sentiment analysis has garnered considerable attention because it incorporates visual data into textual sentiment analysis. Moreover, most current visual-textual sentiment analysis approaches underperform because of their limited exploitation of the correlations between these two modalities. Furthermore, current methods for visual analysis tend to focus excessively on extracting image features while neglecting the aesthetic aspects of images. To address these issues, this study introduces a text-oriented transformer with an image aesthetics assessment fusion network mechanism, ter633133_1_En_12_Chaptermed ToTIAN. This approach comprises two main components: aesthetics-oriented visual feature extraction and a text-oriented transformer. It integrates textual information with image aesthetics—which include emotional cues—via a multiattention mechanism, resulting in a comprehensive representation enriched with emotional cues. Extensive experiments on two publicly available datasets confirmed the superior efficacy of the proposed ToTIAN approach compared with the prevalent unimodal and multimodal methods.