<p>In video summarization, four datasets (TvSum, SumMe, OVP, and YouTube) are typically used for training and testing. In this study, we supplement these datasets with novel features based on High Efficiency Video Codec (HEVC) video coding and motion estimation and compensation. Although HEVC coding variables offer valuable information, they are frequently overlooked in deep learning solutions for video analysis. Thus, we introduce a low-level HEVC feature set suitable for dynamic video summarization. Additionally, we supplement the datasets with additional feature vectors based on CNN embeddings of Window-based Accumulated Image Differences with motion estimation and compensation (WAID-MC). We integrate our two proposed feature sets using feature vector fusion and importance score fusion. Furthermore, we enhance the OVP and YouTube datasets by adding ground truth keyshots, importance scores, and user summaries. In the experimental results section, we compare our proposed solutions with four existing approaches, integrating our features and fusion techniques into their codebases. Our findings indicate that in a minimum of three out of the four datasets, the F1 scores achieved by our proposed methodologies are superior to those of existing studies. Moreover, the experimental results predominantly highlight feature vector fusion as superior to importance score fusion. In many instances, the fusion of the WAID-MC features with the HEVC features yields the best F1 scores. In some cases, the utilization of HEVC features alone results in higher F1 scores compared to existing approaches.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic video summarization using handcrafted features to complement publicly available datasets

  • Zahra Solatidehkordi,
  • Tamer Shanableh

摘要

In video summarization, four datasets (TvSum, SumMe, OVP, and YouTube) are typically used for training and testing. In this study, we supplement these datasets with novel features based on High Efficiency Video Codec (HEVC) video coding and motion estimation and compensation. Although HEVC coding variables offer valuable information, they are frequently overlooked in deep learning solutions for video analysis. Thus, we introduce a low-level HEVC feature set suitable for dynamic video summarization. Additionally, we supplement the datasets with additional feature vectors based on CNN embeddings of Window-based Accumulated Image Differences with motion estimation and compensation (WAID-MC). We integrate our two proposed feature sets using feature vector fusion and importance score fusion. Furthermore, we enhance the OVP and YouTube datasets by adding ground truth keyshots, importance scores, and user summaries. In the experimental results section, we compare our proposed solutions with four existing approaches, integrating our features and fusion techniques into their codebases. Our findings indicate that in a minimum of three out of the four datasets, the F1 scores achieved by our proposed methodologies are superior to those of existing studies. Moreover, the experimental results predominantly highlight feature vector fusion as superior to importance score fusion. In many instances, the fusion of the WAID-MC features with the HEVC features yields the best F1 scores. In some cases, the utilization of HEVC features alone results in higher F1 scores compared to existing approaches.