Video corpus comprises multimedia information located in multiple modalities, which brings evident challenge to emotion-cause extraction tasks. Traditional approaches perform indifferently on multimodal datasets due to the inadequate inter-modal alignment and fusion. We propose a Multi-modal Deep Emotion-Cause Pair Extraction algorithm, which introduces a novel inter-modal attention mechanism to effectively align and fuse text, audio, and video features. The proposed two-stage algorithm first extracts modal pairwise fusion features respectively for emotions and causes. Then, the pairwise features are jointly fused by the emotion-cause fusion module to mine the relationship between emotions and causes. Finally, we utilize a multi-modal fusion module and classifier to identify the emotion-pair relationship among utterances. Experiments show that the proposed algorithm improves the performance of the extraction of multi-modal emotion-cause pairs compared to baseline approaches on the current publicly available dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-modal Deep Emotion-Cause Pair Extraction for Video Corpus

  • Qianli Zhao,
  • Linlin Zong,
  • Bo Xu,
  • Xianchao Zhang,
  • Xinyue Liu

摘要

Video corpus comprises multimedia information located in multiple modalities, which brings evident challenge to emotion-cause extraction tasks. Traditional approaches perform indifferently on multimodal datasets due to the inadequate inter-modal alignment and fusion. We propose a Multi-modal Deep Emotion-Cause Pair Extraction algorithm, which introduces a novel inter-modal attention mechanism to effectively align and fuse text, audio, and video features. The proposed two-stage algorithm first extracts modal pairwise fusion features respectively for emotions and causes. Then, the pairwise features are jointly fused by the emotion-cause fusion module to mine the relationship between emotions and causes. Finally, we utilize a multi-modal fusion module and classifier to identify the emotion-pair relationship among utterances. Experiments show that the proposed algorithm improves the performance of the extraction of multi-modal emotion-cause pairs compared to baseline approaches on the current publicly available dataset.