<p>Video summarization aims to present the most relevant and important information in the video stream in the form of a summary. Most existing researches focus on the selection process of keyframes, determining the importance of video frames by obtaining dependency information between them. However, these works overlook the feature extraction process of video frames. In fact, rich and reliable video frame features are an important basis for determining whether video frames can be selected correctly. This article proposes a new video summarization extraction model based on cross-modal feature fusion (CMFF_VS). CMFF_VS model utilizes the mutual enhancement of video modality and text modality to extract richer semantic information of video frames, thereby providing necessary features for the subsequent video frame selection process. To solve the alignment problem between semantic information of two modalities, CMFF_VS introduces a cross-modal attention mechanism, which utilizes the semantic correlation of modalities to achieve cross-modal semantic fusion. At the same time, CMFF_VS introduces the ASPP module to extract and fuse multi-scale semantic features of individual modalities, enriching the capture of advanced semantic information for each modality. The experimental results show that compared with the state-of-the-art unimodal and multimodal video summarization models, CMFF-VS achieves the best performance, indicating that the cross-modal deep feature extraction and fusion strategy proposed in CMFF-VS is reasonable and effective.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CMFF_VS: A Video Summarization Extraction Model based on Cross-modal Feature Fusion

  • Chaoqun Xin,
  • Mingyang Wang,
  • Xianhao Zhao

摘要

Video summarization aims to present the most relevant and important information in the video stream in the form of a summary. Most existing researches focus on the selection process of keyframes, determining the importance of video frames by obtaining dependency information between them. However, these works overlook the feature extraction process of video frames. In fact, rich and reliable video frame features are an important basis for determining whether video frames can be selected correctly. This article proposes a new video summarization extraction model based on cross-modal feature fusion (CMFF_VS). CMFF_VS model utilizes the mutual enhancement of video modality and text modality to extract richer semantic information of video frames, thereby providing necessary features for the subsequent video frame selection process. To solve the alignment problem between semantic information of two modalities, CMFF_VS introduces a cross-modal attention mechanism, which utilizes the semantic correlation of modalities to achieve cross-modal semantic fusion. At the same time, CMFF_VS introduces the ASPP module to extract and fuse multi-scale semantic features of individual modalities, enriching the capture of advanced semantic information for each modality. The experimental results show that compared with the state-of-the-art unimodal and multimodal video summarization models, CMFF-VS achieves the best performance, indicating that the cross-modal deep feature extraction and fusion strategy proposed in CMFF-VS is reasonable and effective.