DeepK-Means: A Fusion of DCNN and K-Means Clustering for Video Summarization
摘要
Video summarization techniques aim to distill the most salient content from videos, constructing a brief yet comprehensive synopsis. Over the past two decades, numerous strategies have emerged, with recent deep neural network (DNN) based methods setting the benchmark in performance and reliability. In this study, we introduce a novel video summarization method leveraging Deep Convolutional Neural Networks and k-means clustering, which we denote as “DeepK-means.” Our primary contribution lies in utilizing DeepK-means to extract pivotal video highlights efficiently. We detail the salient features adopted to craft these video summaries, emphasizing our method’s proficiency in detecting key highlights. Our experiments employ the TVSum50 dataset, allowing for a rigorous evaluation of our technique. Further, we furnish performance metrics to objectively assess video summarization algorithms and juxtapose the effectiveness of our proposed DeepK-means approach against a contemporary deep learning technique, focusing on precision. Impressively, DeepK-means achieved a precision of 99.64%, with recall and f1-score closely following at 99.63%. Based on these promising results, we suggest potential avenues for future research, touching upon topics like the scalability of annotated data and the broader utility of our evaluation metrics.