Text-to-image person re-identification aims to identify a target person from a large-scale image gallery based on a given textual query. The primary challenge in this task lies in bridging the semantic gap between images and texts. Current methods align texts and images into a joint embedding space. Nevertheless, these approaches overlook the inter-modal information inconsistency and the intra-modal semantic inconsistency, leading to the ineffective construction of the intricate semantic correspondence between texts and images. To address the above challenges, we propose a text-to-image person re-identification method based on cross-modal dual matching and comparison. Specifically, we propose an adaptive dual local matching module. This module predicts masked textual tokens with the image patches most relevant to the texts, and reconstructs masked image patches with the texts and their least relevant image patches. These processes ensure the model focuses on the body parts biased by the text descriptions, thereby implicitly establishing fine-grained associations of cross-modal semantic. Furthermore, we propose a cross-modal dual contrastive loss that performs inter-modal and intra-modal contrastive learning to minimize the distance between matched image-text pairs and ensure the distance between different matched image-text pairs, thereby maintaining inter-modal and intra-modal semantic consistency in a joint embedding space. The experimental results demonstrate that our method performs comparably to or even better than existing methods, obtaining Rank-1 of 74.68% on the CUHK-PEDES dataset and 65.29% on the ICFG-PEDES dataset, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Modal Dual Matching and Comparison for Text-to-Image Person Re-identification

  • Lin Cao,
  • Wenwen Sun,
  • Yanan Guo,
  • Shoujing Wang,
  • Boqian Lv

摘要

Text-to-image person re-identification aims to identify a target person from a large-scale image gallery based on a given textual query. The primary challenge in this task lies in bridging the semantic gap between images and texts. Current methods align texts and images into a joint embedding space. Nevertheless, these approaches overlook the inter-modal information inconsistency and the intra-modal semantic inconsistency, leading to the ineffective construction of the intricate semantic correspondence between texts and images. To address the above challenges, we propose a text-to-image person re-identification method based on cross-modal dual matching and comparison. Specifically, we propose an adaptive dual local matching module. This module predicts masked textual tokens with the image patches most relevant to the texts, and reconstructs masked image patches with the texts and their least relevant image patches. These processes ensure the model focuses on the body parts biased by the text descriptions, thereby implicitly establishing fine-grained associations of cross-modal semantic. Furthermore, we propose a cross-modal dual contrastive loss that performs inter-modal and intra-modal contrastive learning to minimize the distance between matched image-text pairs and ensure the distance between different matched image-text pairs, thereby maintaining inter-modal and intra-modal semantic consistency in a joint embedding space. The experimental results demonstrate that our method performs comparably to or even better than existing methods, obtaining Rank-1 of 74.68% on the CUHK-PEDES dataset and 65.29% on the ICFG-PEDES dataset, respectively.