错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Appearance-Motion Dual-Stream Heterogeneous Network for VideoQA

  • Feifei Xu,
  • Zheng Zhong,
  • Yitao Zhu,
  • Yingchen Zhou,
  • Guangzhen Li

摘要

Capturing spatio-temporal information in videos related to the question remains a key challenge in video question answering task (VideoQA). Though great success has been achieved in VideoQA, most of the existing methods do not sufficiently consider the correlation among appearance, motion, and object features, making it difficult to fully exploit the spatio-temporal relationships at different granularities. Besides, recent researches typically use the same interaction method when different features in the video interact with the question features separately, which ignores the spatio-temporal characteristics of the appearance and motion features in the video which leads to the problem of spatio-temporal mismatch. In this paper, we propose an Appearance-Motion Dual-stream Heterogeneous Network for VideoQA (AMHN), which pays attention to the synergy among three different features by heterogeneous interactions in terms of their spatio-temporal characteristics. AMHN unites object features with appearance features and motion features respectively to obtain two high-level visual representations containing object information. Then they are fed into the object-relational reasoning module to acquire relation-aware visual features. We use a bilinear attention network for appearance and put forward a Video-Text Symmetric Attention Network (VTSAN) for motion to achieve diverse features, which are fused under the guidance of the question to predict the final answer. We evaluate the performance of AMHN on two VideoQA benchmark datasets and perform an extensive ablation study. The experimental results demonstrate its state-of-the-art.