<p>As a fundamental and challenging task, image-text retrieval has attracted significant attention in recent years. The primary challenge lies in bridging the modality gap between image and text to achieve precise retrieval. While existing works focus on learning common embedding or cross-modal alignment for image-text data pairs, they mainly utilize information solely from the data pairs themselves. However, the potential benefits of leveraging implicit and explicit knowledge for bridging the modality gap have not been well explored. For this purpose, we propose a novel Implicit and Explicit Knowledge Enhanced cross-modal Representation (IEKER) network for image-text retrieval. Instead of directly learning common feature representation or aligning image and text through cross-modal interaction, we adopt two types of knowledge as intermediaries to enhance the heterogeneous modalities representation. The first type is the implicit knowledge in form of a modality-common memory bank of sharing features learned from image and text modalities. The second type is the explicit knowledge formatted as knowledge graph constructed from the high-frequency semantic concepts or labels of image captions. Through extensive experiments on the Flicker30K and MS-COCO datasets, the compelling performance of the proposed IEKER network validates its effectiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Implicit and explicit knowledge enhanced cross-modal representation for image-text retrieval

  • Hua Huang,
  • Qiaoli Qin

摘要

As a fundamental and challenging task, image-text retrieval has attracted significant attention in recent years. The primary challenge lies in bridging the modality gap between image and text to achieve precise retrieval. While existing works focus on learning common embedding or cross-modal alignment for image-text data pairs, they mainly utilize information solely from the data pairs themselves. However, the potential benefits of leveraging implicit and explicit knowledge for bridging the modality gap have not been well explored. For this purpose, we propose a novel Implicit and Explicit Knowledge Enhanced cross-modal Representation (IEKER) network for image-text retrieval. Instead of directly learning common feature representation or aligning image and text through cross-modal interaction, we adopt two types of knowledge as intermediaries to enhance the heterogeneous modalities representation. The first type is the implicit knowledge in form of a modality-common memory bank of sharing features learned from image and text modalities. The second type is the explicit knowledge formatted as knowledge graph constructed from the high-frequency semantic concepts or labels of image captions. Through extensive experiments on the Flicker30K and MS-COCO datasets, the compelling performance of the proposed IEKER network validates its effectiveness.