Multimodal Local Feature Enhancement Network for Video Summarization
摘要
Multimodal information processing has garnered considerable attention in recent years. Due to the inherent multimodal information in videos, multimodal learning has been introduced in the domain of video summarization, leading to a significant improvement in summary accuracy. However, previous multimodal video summarization methods have overlooked the local fine-grained features of the fused multimodal information, constraining the representational capacity of the models. Therefore, we propose a Multimodal Local Feature Enhancement Network (MLFEN) to address this issue. Firstly, we utilize multi-head attention to encode single-modal information, capturing their mutual dependencies, and then fuse multiple modalities. Secondly, we further process the multimodal information by employing local attention to explore local fine-grained interdependencies and global attention to capture global dependencies. Finally, a regression network is employed to predict the frame-level importance scores used in generating video summaries. In addition, we explored different fusion methods for integrating different modalities of information. Our method demonstrates superior performance on both evaluation metrics, achieving F-measure of 0.528, Kendall’s \(\tau \) of 0.076, and Spearman’s \(\rho \) of 0.099 on the SumMe dataset, as well as F-measure of 0.612, Kendall’s \(\tau \) of 0.155, and Spearman’s \(\rho \) of 0.204 on the TVSum dataset. These results surpass the performance of most existing video summarization methods.