VASEC: Visual Attention-Driven Key Shot Extraction and Video Enhancement with Parallelly Decoded Dense Captions
摘要
In today’s information era, vast amount of multimedia content has surged the necessity for techniques to detect significant and relevant content. We introduce visual attention-driven model (VAFF) for extracting the prominent shots of a video which is followed by enhancement and event-wise video captioning where parallel computation is exploited. The pre-trained GoogleNet model is utilized to obtain the visual features (object/static) used for key shot extraction and pre-trained I3D model is utilized to obtain the kinetic features (RGB and Flow). The modified version of 0/1 Knapsack Algorithm used the predicted importance score for key shot extraction. F1 score for the Visual Attention Feature Fusion (VAFF) showed improvement for SumMe and considerable value for TVSum due to their respective lengths. The shortened video is enhanced using pre-trained ESRGAN. Dense video captions are generated using the Temporally Sensitive Pre-trained model (TSP) features extracted from the enhanced video. Thus, research aims to provide a robust pipeline for static and real-time video analysis.