Fine-grained Feature Assisted Cross-modal Image-text Retrieval
摘要
Cross-modal image-text retrieval is a challenging task due to the inherent ambiguity between modalities. However, most existing methods formulate this problem either with the coarse-grained information of the global image, ignoring the valuable fine-grained information implicit in local instances, or with the local features of the images and words, failing to provide a global understanding. In this paper, we propose a novel Fine-grained Feature Assisted Cross-modal Image-Text Retrieval (FiACR) model to learn a comprehensive and informative visual representation for cross-modal retrieval. Specifically, to address the absence of local information, we design a Local-Global Visual Features Fusion (LGVFF) module to aggregate global image and local instance information. By aggregation, FiACR can capture and leverage the images’ intricate details, which enables an accurate alignment between image and text. To enhance the global visual representation capability, we utilize the instance features to filter the global image feature’s attention and encourage it to focus on prominent regions in the image. Experimental results on several datasets show the competitive accuracy of our method compared to prior art.