错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Local Feature Enhancement Network for Video Summarization

  • Zhaoyun Li,
  • Xiwei Ren,
  • Fengyi Du

摘要

Multimodal information processing has garnered considerable attention in recent years. Due to the inherent multimodal information in videos, multimodal learning has been introduced in the domain of video summarization, leading to a significant improvement in summary accuracy. However, previous multimodal video summarization methods have overlooked the local fine-grained features of the fused multimodal information, constraining the representational capacity of the models. Therefore, we propose a Multimodal Local Feature Enhancement Network (MLFEN) to address this issue. Firstly, we utilize multi-head attention to encode single-modal information, capturing their mutual dependencies, and then fuse multiple modalities. Secondly, we further process the multimodal information by employing local attention to explore local fine-grained interdependencies and global attention to capture global dependencies. Finally, a regression network is employed to predict the frame-level importance scores used in generating video summaries. In addition, we explored different fusion methods for integrating different modalities of information. Our method demonstrates superior performance on both evaluation metrics, achieving F-measure of 0.528, Kendall’s \(\tau \) of 0.076, and Spearman’s \(\rho \) of 0.099 on the SumMe dataset, as well as F-measure of 0.612, Kendall’s \(\tau \) of 0.155, and Spearman’s \(\rho \) of 0.204 on the TVSum dataset. These results surpass the performance of most existing video summarization methods.