<p>Multimodal sentiment analysis (MSA) integrates and processes data from multiple sources, like audio and text, to better understand human emotions through cross-modal interactions. The effective acquisition and integration of meaningful features for constructing richer sentiment representations remains a key challenge in MSA. Most existing methods directly obtain global representations and integrate at the utterance level from different modalities, but this ignores fine-grained representations and makes it difficult to capture intricate relationships within and between modalities. Therefore, we propose a novel method, Fine-grained Multimodal Fusion Network (MMTA). Firstly, a Fine-grained Alignment (FGA) module is introduced to align and extract word-level features to bridge heterogeneous modal gaps. FGA enables word-level alignment between audio, text, and their corresponding contextual information using the Montreal Forced Aligner (MFA). Secondly, a Multi-level Fusion module (MLF) is designed, which captures more cross-modal interaction through three stages: Local-Local Interaction, Local-Global Interaction, and Similarity-weighted Representation Adjustment. Finally, an Attention Fusion Network(AFN) module is developed to capture both inter- and intra-modal correlations, enabling the generation of consistent multimodal representations. Extensive evaluations on widely used MSA datasets, CMU-MOSI and CMU-MOSEI, indicate that our method outperforms prior baselines and validates the effectiveness of the fine-grained alignment and the multi-level fusion for improving multimodal sentiment analysis performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-level fusion with fine-grained alignment for multimodal sentiment analysis

  • Xiaoge Li,
  • Yanan Ma,
  • Xiaochun An,
  • Jinshuo Xing,
  • Ren Liu,
  • Yunsheng Ren

摘要

Multimodal sentiment analysis (MSA) integrates and processes data from multiple sources, like audio and text, to better understand human emotions through cross-modal interactions. The effective acquisition and integration of meaningful features for constructing richer sentiment representations remains a key challenge in MSA. Most existing methods directly obtain global representations and integrate at the utterance level from different modalities, but this ignores fine-grained representations and makes it difficult to capture intricate relationships within and between modalities. Therefore, we propose a novel method, Fine-grained Multimodal Fusion Network (MMTA). Firstly, a Fine-grained Alignment (FGA) module is introduced to align and extract word-level features to bridge heterogeneous modal gaps. FGA enables word-level alignment between audio, text, and their corresponding contextual information using the Montreal Forced Aligner (MFA). Secondly, a Multi-level Fusion module (MLF) is designed, which captures more cross-modal interaction through three stages: Local-Local Interaction, Local-Global Interaction, and Similarity-weighted Representation Adjustment. Finally, an Attention Fusion Network(AFN) module is developed to capture both inter- and intra-modal correlations, enabling the generation of consistent multimodal representations. Extensive evaluations on widely used MSA datasets, CMU-MOSI and CMU-MOSEI, indicate that our method outperforms prior baselines and validates the effectiveness of the fine-grained alignment and the multi-level fusion for improving multimodal sentiment analysis performance.