Two Uni-directional LSTMs-Based Captioning Module for Dense Video Captioning
摘要
Most real-life videos encompass a diverse array of events, which may be sequential or overlapping. Dense video captioning involves the identification and localization of events within a video, as well as the generation of descriptive captions for each event. To accurately describe an event at a specific time step, it is crucial to integrate contextual information from both past and future events. Our proposed approach utilizes two uni-directional LSTM-based captioning modules that synthesize contextual information from both visual and textual data in forward and backward directions to generate dense video caption. This model demonstrates a significant advancement over the leading dense video captioning methods, achieving a relative improvement of approximately 9% on the ActivityNet dataset.