<p>Multimodal sentiment analysis benefits from leveraging text, video, and audio data to better infer emotional states, but the heterogeneous nature of these modalities presents challenges. This paper proposes the adaptive text-guided multimodal gated-fusion transformer (ATMGT), a novel model designed to address the challenges. The ATMGT leverages Transformer-based self-attention and cross-attention mechanisms to enable the deep interaction and fusion of the text, audio, and visual modalities. Then, a gated fusion mechanism effectively integrates these audio-visual features, alleviating information redundancy. Additionally, text is treated as a local feature to guide the scaling of global features formed by audio-visual data, emphasizing emotionally significant regions. Furthermore, a self-supervised label generation module enhances modality-specific learning and robust sentiment classification. Experimental comparisons demonstrate that the proposed model achieves competitive performance across multiple metrics on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets, outperforming state-of-the-art models. Finally, ablation studies validate the contributions of the core modules to the overall model performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-Guided Enhanced Transformer Fusion for Multimodal Sentiment Analysis

  • Xiangyu Liao,
  • Xianxin Ke,
  • Qinghua Liu,
  • Guoliang Li

摘要

Multimodal sentiment analysis benefits from leveraging text, video, and audio data to better infer emotional states, but the heterogeneous nature of these modalities presents challenges. This paper proposes the adaptive text-guided multimodal gated-fusion transformer (ATMGT), a novel model designed to address the challenges. The ATMGT leverages Transformer-based self-attention and cross-attention mechanisms to enable the deep interaction and fusion of the text, audio, and visual modalities. Then, a gated fusion mechanism effectively integrates these audio-visual features, alleviating information redundancy. Additionally, text is treated as a local feature to guide the scaling of global features formed by audio-visual data, emphasizing emotionally significant regions. Furthermore, a self-supervised label generation module enhances modality-specific learning and robust sentiment classification. Experimental comparisons demonstrate that the proposed model achieves competitive performance across multiple metrics on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets, outperforming state-of-the-art models. Finally, ablation studies validate the contributions of the core modules to the overall model performance.