To address the limitations of the existing DSNet model in video summarization, including insufficient temporal information utilization, poor multi-modal fusion, and weak multi-scale feature capture capabilities leading to incoherent contexts and low-quality summaries, this paper proposes a video summarization algorithm based on multi-modal multi-scale temporal conjugate positional encoding. The framework integrates temporal conjugate positional encoding into the multi-head attention mechanism to fuse video, audio, and text modalities, while introducing a multi-scale feature pyramid structure for video segmentation. The algorithm uses temporal encoding and multi-modal fusion for temporal modeling, boosts cross-hierarchical features via the pyramid, and hierarchically predicts important frame and video segment. Combined with segment classification-regression methods, it simultaneously outputs shot confidence scores and positional offsets to select key shots. Experimental results on standard and augmented versions of SumMe and TVSum datasets demonstrate F-Score values of 54.5%, 54.9% and 62.6%, 63.0%, outperforming the VJMHT model by 3.2%, 2.7% and 1.7%, 1.1% respectively. This demonstrates the algorithm’s enhanced summarization accuracy and theoretical contributions to video summarization.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Video Summarization Algorithm Based on Multimodal Multiscale Temporal Conjugate Position Coding

  • Teng Liu,
  • Yude Wang,
  • Mingtian Wang,
  • Xinyu Wang,
  • Haoyu Guo,
  • Yuanpei Wang,
  • Fei Song,
  • Ting Zhao

摘要

To address the limitations of the existing DSNet model in video summarization, including insufficient temporal information utilization, poor multi-modal fusion, and weak multi-scale feature capture capabilities leading to incoherent contexts and low-quality summaries, this paper proposes a video summarization algorithm based on multi-modal multi-scale temporal conjugate positional encoding. The framework integrates temporal conjugate positional encoding into the multi-head attention mechanism to fuse video, audio, and text modalities, while introducing a multi-scale feature pyramid structure for video segmentation. The algorithm uses temporal encoding and multi-modal fusion for temporal modeling, boosts cross-hierarchical features via the pyramid, and hierarchically predicts important frame and video segment. Combined with segment classification-regression methods, it simultaneously outputs shot confidence scores and positional offsets to select key shots. Experimental results on standard and augmented versions of SumMe and TVSum datasets demonstrate F-Score values of 54.5%, 54.9% and 62.6%, 63.0%, outperforming the VJMHT model by 3.2%, 2.7% and 1.7%, 1.1% respectively. This demonstrates the algorithm’s enhanced summarization accuracy and theoretical contributions to video summarization.