Learning multimodal features of videos is a key for accurately understanding the real emotion expressed in the videos. Sequence alignment is usually necessary to deal with different modal sequence lengths in order to fuse multimodal features together. Text Enhancement-based Multimodal Fusion (TEMF) is implemented in the paper to integrate multimodal information without sequence alignment for improving video sentiment analysis. For modeling the unaligned sequences of multimodal inputs, text is enhanced by cross-modal attention mechanism. The expression level of non-textual features is strengthened so that textual and non-textual features interact at similar representational level. Experiments are conducted with three benchmark datasets and demonstrate TEMF outperforms several comparison methods in terms of Acc and F1.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text Enhancement-Based Multimodal Fusion for Video Sentiment Analysis

  • Junhao Luo,
  • Yan Zhu,
  • Yiqiang Peng

摘要

Learning multimodal features of videos is a key for accurately understanding the real emotion expressed in the videos. Sequence alignment is usually necessary to deal with different modal sequence lengths in order to fuse multimodal features together. Text Enhancement-based Multimodal Fusion (TEMF) is implemented in the paper to integrate multimodal information without sequence alignment for improving video sentiment analysis. For modeling the unaligned sequences of multimodal inputs, text is enhanced by cross-modal attention mechanism. The expression level of non-textual features is strengthened so that textual and non-textual features interact at similar representational level. Experiments are conducted with three benchmark datasets and demonstrate TEMF outperforms several comparison methods in terms of Acc and F1.