Element-Centered Multi-granularity Network for Dense Video Captioning
摘要
With the development of deep learning technologies, the task of video captioning has been extended from single-sentence description for short and trimmed videos to multi-sentence description for long and untrimmed videos. Existing works concentrate mainly on how to improve the relevance and coherence of generated descriptions, based on criteria such as BLEU, METEOR, and CIDEr. While used for purposes such as assisting disabled people, authenticity must be considered prior to coherence and other criteria. Besides, unlike other tasks in video understanding which always have an objective and only solution to provide, such as video action classification, video event localization and etc., descriptions provided by each individual for the same video can vary from the definition of video events, to the sentence structure and the exact words one chooses. Such an one-to-many correspondence relationship in video-text pairs make it unavoidable to have inconsistent labels for the same video, and thus poses a challenge to the learning ability of captioning models. In order to alleviate the effect of inconsistency in labels and improve the authenticity of generated captions, we focus on the elements which are necessary for both human and models to define an event, and based on which we propose a novel Element-Centered Multi-Granularity Network (ECMGN) for dense video captioning, which can utilize labels for different purposes and concentrate on the core elements of video events to generate more reliable descriptions for given videos. Experimental results demonstrate that our model significantly improves the quality of generated descriptions both in traditional and authenticity metrics.