Accurate emotion detection relies on the fusion of information from multiple modalities, including text, audio, and video. However, existing fusion methods face challenges such as information redundancy and significant differences in representation between modalities, making it difficult to detect subtle emotional cues, particularly in audio and visual data, and when unimodal annotations are unavailable. In response to the aforementioned issues, we propose the Fine-Grained Multimodal Fusion Network (FG-MFN), which effectively captures subtle emotional variations by adopting multi-directional and multi-scale feature extraction techniques. Specifically, we introduce a fine-grained fusion mechanism that enhances feature representations by promoting cross-dimensional interaction through rotation operations and residual transformations. Additionally, a self-supervised mechanism is employed to generate modality-wise labels, reducing annotation costs and enhancing modality-specific learning. Numerous experiments on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets show that FG-MFN consistently surpasses the leading baselines. Quantitative and qualitative results validate the robustness of the proposed approach across multilingual settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-Grained Feature Fusion for Self-supervised Multimodal Sentiment Analysis

  • Tan Deng,
  • Shiyu Mei,
  • Hujin Peng,
  • Xingyu Du,
  • Zeyu Chen,
  • Mingfeng Huang,
  • Xiaoyong Tang,
  • Ronghui Cao,
  • Wenzheng Liu

摘要

Accurate emotion detection relies on the fusion of information from multiple modalities, including text, audio, and video. However, existing fusion methods face challenges such as information redundancy and significant differences in representation between modalities, making it difficult to detect subtle emotional cues, particularly in audio and visual data, and when unimodal annotations are unavailable. In response to the aforementioned issues, we propose the Fine-Grained Multimodal Fusion Network (FG-MFN), which effectively captures subtle emotional variations by adopting multi-directional and multi-scale feature extraction techniques. Specifically, we introduce a fine-grained fusion mechanism that enhances feature representations by promoting cross-dimensional interaction through rotation operations and residual transformations. Additionally, a self-supervised mechanism is employed to generate modality-wise labels, reducing annotation costs and enhancing modality-specific learning. Numerous experiments on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets show that FG-MFN consistently surpasses the leading baselines. Quantitative and qualitative results validate the robustness of the proposed approach across multilingual settings.