<p>Recent advances in sentiment analysis have primarily focused on fusing multimodal information from video data, including visual, acoustic, and textual features, across temporal sequences. While great effort has been made to integrate or fuse information across modalities, less is known about the extent to which temporal segments contribute to model decisions. In addition, current interpretable methods, such as prototype networks, are primarily designed for uni-modal analysis and fail to handle the complex interactions between multiple modalities and temporal dependencies inherent in video data. To address the challenges, we propose <b>M</b>ulti<b>M</b>odal <b>P</b>rototypical <b>Net</b>works (MMPNet), which extends prototype-based interpretability to multimodal sentiment classification. Specifically, MMPNet can identify contributions of time-level features and leverage them to explain why a particular prediction was made, while also helping to find the relative importance of modality-level features. Experimental results show that MMPNet outperforms existing methods by 2.9% and 1.6% in accuracy on CMU-MOSI and CMU-MOSEI respectively, and achieves better interpretability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal prototypical network for interpretable sentiment classification

  • Chenguang Song,
  • Ke Chao,
  • Bingjing Jia,
  • Yiqing Shen

摘要

Recent advances in sentiment analysis have primarily focused on fusing multimodal information from video data, including visual, acoustic, and textual features, across temporal sequences. While great effort has been made to integrate or fuse information across modalities, less is known about the extent to which temporal segments contribute to model decisions. In addition, current interpretable methods, such as prototype networks, are primarily designed for uni-modal analysis and fail to handle the complex interactions between multiple modalities and temporal dependencies inherent in video data. To address the challenges, we propose MultiModal Prototypical Networks (MMPNet), which extends prototype-based interpretability to multimodal sentiment classification. Specifically, MMPNet can identify contributions of time-level features and leverage them to explain why a particular prediction was made, while also helping to find the relative importance of modality-level features. Experimental results show that MMPNet outperforms existing methods by 2.9% and 1.6% in accuracy on CMU-MOSI and CMU-MOSEI respectively, and achieves better interpretability.