Text-Image relation inference (TIRI) aims to identify the potential semantic relationships between text and image. Although previous works have made some progress, there are still two issues that are worth exploring. First, existing studies primarily rely on the English text and the accompanying image to perform TIRI. This completely ignores other external information (e.g, another parallel the linguistic knowledge from Chinese by machine translation). As previous studies, bilingual knowledge can help us understand the semantics of the text more accurately, like polysemy problems. Second, existing studies normally employ a Transformer-based structure and implicitly encode different modalities. This completely neglects the potential dependencies among each unit in the uni-modality (e.g., the dependency syntax in the text). Therefore, we propose a bilingual multimodal graph convolutional network (BMGCN) to model both intra-modal and inter-modal dependence in a fine-grained manner. This approach can not only explicitly model the dependencies within each modality and each language, but also model the dependencies between different modalities and different languages simultaneously. Systematic experiments demonstrate that our BMGCN obviously outperforms the state-of-the-art on two datasets. Additionally, we provide several interesting analyses to further verify the effectiveness of our proposed approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bilingual Multimodal Graph Modeling for Text-Image Relation Inference

  • Dong Zhang,
  • Wenjie Lu,
  • Shoushan Li,
  • Guodong Zhou

摘要

Text-Image relation inference (TIRI) aims to identify the potential semantic relationships between text and image. Although previous works have made some progress, there are still two issues that are worth exploring. First, existing studies primarily rely on the English text and the accompanying image to perform TIRI. This completely ignores other external information (e.g, another parallel the linguistic knowledge from Chinese by machine translation). As previous studies, bilingual knowledge can help us understand the semantics of the text more accurately, like polysemy problems. Second, existing studies normally employ a Transformer-based structure and implicitly encode different modalities. This completely neglects the potential dependencies among each unit in the uni-modality (e.g., the dependency syntax in the text). Therefore, we propose a bilingual multimodal graph convolutional network (BMGCN) to model both intra-modal and inter-modal dependence in a fine-grained manner. This approach can not only explicitly model the dependencies within each modality and each language, but also model the dependencies between different modalities and different languages simultaneously. Systematic experiments demonstrate that our BMGCN obviously outperforms the state-of-the-art on two datasets. Additionally, we provide several interesting analyses to further verify the effectiveness of our proposed approach.