<p>Multimodal Sentiment Analysis is known for its ability to more comprehensively predict the emotional tendencies of users. The data fusion module plays a crucial role in multimodal sentiment analysis, enabling the integration of information from multiple modalities. However, effectively fusing these modal data is a challenging task. Some methods using CBAM technology have shown significant performance improvements in computer vision and deep learning tasks. Inspired by the inherent structure of CBAM-based models, we have discovered that it can be naturally applied to multimodal feature processing. To this end, we propose TriAxial Modality Attention Fusion (TAMAttention). TAMAttention accepts all relevant modal features as input and mixes them in three dimensions: sequential (L), modality (M), and channel (D). During the mixing process, different multimodal information is effectively transferred and shared to extract important features related to emotional tendencies. In addition, we use a top-down fusion mechanism to fully exploit the correlation and complementarity between different modalities to improve the performance of sentiment analysis tasks. We conducted in-depth experiments on two sentiment analysis datasets, CMU-MOSI and CMU-MOSEI, and the results show that TAMAttention has made significant progress in predicting sentiment tendencies and is highly competitive with existing sentiment analysis methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Triaxial modality attention fusion with top-down mask generation for enhanced multimodal sentiment analysis

  • Cheng Feng,
  • Hai Yang,
  • Shuxian Wang,
  • Xue Li

摘要

Multimodal Sentiment Analysis is known for its ability to more comprehensively predict the emotional tendencies of users. The data fusion module plays a crucial role in multimodal sentiment analysis, enabling the integration of information from multiple modalities. However, effectively fusing these modal data is a challenging task. Some methods using CBAM technology have shown significant performance improvements in computer vision and deep learning tasks. Inspired by the inherent structure of CBAM-based models, we have discovered that it can be naturally applied to multimodal feature processing. To this end, we propose TriAxial Modality Attention Fusion (TAMAttention). TAMAttention accepts all relevant modal features as input and mixes them in three dimensions: sequential (L), modality (M), and channel (D). During the mixing process, different multimodal information is effectively transferred and shared to extract important features related to emotional tendencies. In addition, we use a top-down fusion mechanism to fully exploit the correlation and complementarity between different modalities to improve the performance of sentiment analysis tasks. We conducted in-depth experiments on two sentiment analysis datasets, CMU-MOSI and CMU-MOSEI, and the results show that TAMAttention has made significant progress in predicting sentiment tendencies and is highly competitive with existing sentiment analysis methods.