Dense Video Captioning with Context Fusion and Reasoning
摘要
Dense video captioning represents a challenge within multimedia processing, designed to detect and depict key events in the video through a series of descriptive sentences. However, earlier researches for dense video captioning mainly relied on visual cues within videos, falling short in accurately identifying and detailing overlapping and/or long-lasting events. We propose a dense video captioning method named Bidirectional Relational Recurrent Neural Network (Bi-RRNN), that leverages both local and global contextual information, along with visual information from video. Specifically, a bidirectional recurrent neural network is employed to capture the global and local context surrounding a specific event, and visual information is extracted from the video using a 3D convolutional neural network. Thus, Bi-RRNN could capture events within a video, reason/ratiocinate the relations between these events, produce a series of descriptive sentences that accurately reflect the video’s content. Experimental results show that Bi-RRNN perform well on the ActivityNet Captions dataset.