Numerical and spatiotemporal features fusion for video summarization
摘要
In this paper, we propose a numerical spatiotemporal approach for video summarization. The current solutions leverage deep learning techniques to tackle this task. However, existing methods do not employ the video shots’ length data in their networks. We first introduce a new ground truth labelling for video summarization. This ground truth tackles multiple users’ annotations and is inclusive of video shot length information. We then propose a novel network architecture Numerical SpatioTemporal Fusion Net, NSTFNet. The proposed architecture leverages the temporal modelling ability of the self-attention mechanism to model video’s vision data to be fused with structured numerical features of video shots. We evaluate our results on TVSum and SumMe datasets. Experimental results show that the proposed model outperforms state-of-the-art performance on the SumMe dataset with 51.29% and 54.42% f-scores on the canonical and augmented configurations, demonstrating the effectiveness of our proposed model compared to state-of-the-art methods.