错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph-based relational reasoning network for video question answering

  • Tao Tan,
  • Guanglu Sun

摘要

Vdeo quesition answering models need to accurately understand video content based on the question and reason out the correct answer. The relations between objects in the video can enhance the ability of the model to comprehend video information. However, due to the complex stereoscopic spatio-temporal structure among objects in the video, current methods still lack sufficient capability to reason about the dynamic spatio-temporal relations between objects. Therefore, this paper proposes a graph-based relational reasoning network for video question answering. For dynamic spatio-temporal relations, a stereoscopic spatio-temporal graph is designed that models the stereoscopic spatio-temporal structure among objects in the video. Spatio-temporal graph convolutional network is employed to synchronously reason about the static spatial relations among objects within the same frame and the dynamic temporal relations among objects in adjacent frames. For appearance relations, globally-aware object-level appearance features are used to construct a fully connected appearance graph. Appearance graph convolutional network is employed to reason about the appearance relations among objects. Our model is evaluated on benchmark datasets and subjected to extensive ablation studies. The experimental results demonstrate the effectiveness of the model presented in this paper.