错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text Visual Question Answering Based on Interactive Learning and Relationship Modeling

  • Chao Zhang,
  • Wei Wu,
  • Bingzhuo Ma

摘要

The Text Visual Question Answering (TextVQA) task is introduced to understand textual information and answer the question related to textual information in daily life scenarios. In the TextVQA task, the interaction fusion between different modalities (question text, visual objects, and visual text) plays an important role, which can capture the correlation between different modalities, thereby improving ability to understand and answer the question accurately. However, most existing methods cannot effectively handle the interaction fusion between modalities well. Therefore, in order to more effectively integrate information from different modalities, this paper proposes a module called Multimodal Feature Cascade Guidance (MFCG), which solves the problem of ignoring the importance of certain modalities in previous methods. In addition, a Relative Position Relationship Enhanced Transformer (RPRET) layer is introduced to model the relative position relationship between different modalities in the image, thereby improving the performance of answering the question related to spatial position relationships. The proposed method outperforms various state-of-the-art models on two public datasets, which confirms the effectiveness of our method.