<p>Video-based commonsense captioning is a core task in the field of video understanding that aims to generate underlying commonsense knowledge. Existing studies leverage rich visual and text information for multimodel video captioning, however, neglect the fine-grained spatial and temporal features in videos. To address these issues, this paper proposes STAGVid2C, a novel model that utilizes the Spatio-Temporal Action Graph to enhance Video-based Commonsense Captioning. The graph uses the grid features of video frames as nodes, and employs the spatial connections and temporal similarity between video frames as edges. Then, a graph neural network is employed to learn the graph representation which can capture more semantic relationships between image and action in videos. After that, STAGVid2C uses a memory network to facilitate cross-modal fusion and alignment, thereby generating more accurate commonsense captions. Experimental results on the public V2C dataset demonstrate that our proposed STAGVid2C significantly outperforms the state-of-the-art methods. Especially, in terms of CIDEr, it shows an average score improvement of 7.3% in the intention part and 8.7% in the effect part.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

STAGVid2C: enhancing video-based commonsense captioning with spatio-temporal action graph

  • Haitao Xiong,
  • Junhong Ding,
  • Yuchen Zhou,
  • Yuanyuan Cai

摘要

Video-based commonsense captioning is a core task in the field of video understanding that aims to generate underlying commonsense knowledge. Existing studies leverage rich visual and text information for multimodel video captioning, however, neglect the fine-grained spatial and temporal features in videos. To address these issues, this paper proposes STAGVid2C, a novel model that utilizes the Spatio-Temporal Action Graph to enhance Video-based Commonsense Captioning. The graph uses the grid features of video frames as nodes, and employs the spatial connections and temporal similarity between video frames as edges. Then, a graph neural network is employed to learn the graph representation which can capture more semantic relationships between image and action in videos. After that, STAGVid2C uses a memory network to facilitate cross-modal fusion and alignment, thereby generating more accurate commonsense captions. Experimental results on the public V2C dataset demonstrate that our proposed STAGVid2C significantly outperforms the state-of-the-art methods. Especially, in terms of CIDEr, it shows an average score improvement of 7.3% in the intention part and 8.7% in the effect part.