错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dense Video Captioning with Context Fusion and Reasoning

  • Wanting Ji,
  • Hao Qin,
  • Ruili Wang,
  • Tingwei Chen

摘要

Dense video captioning represents a challenge within multimedia processing, designed to detect and depict key events in the video through a series of descriptive sentences. However, earlier researches for dense video captioning mainly relied on visual cues within videos, falling short in accurately identifying and detailing overlapping and/or long-lasting events. We propose a dense video captioning method named Bidirectional Relational Recurrent Neural Network (Bi-RRNN), that leverages both local and global contextual information, along with visual information from video. Specifically, a bidirectional recurrent neural network is employed to capture the global and local context surrounding a specific event, and visual information is extracted from the video using a 3D convolutional neural network. Thus, Bi-RRNN could capture events within a video, reason/ratiocinate the relations between these events, produce a series of descriptive sentences that accurately reflect the video’s content. Experimental results show that Bi-RRNN perform well on the ActivityNet Captions dataset.