Text-based person retrieval primarily aims to retrieve the images of target persons represented by a given text query. In this task, how to effectively align images and text globally and locally is an important challenge. At the same time, since multiple modalities are involved, reducing the differences between different modalities is also an important challenge. Existing work has focused more on reconstructing an image rather than the semantic consistency between the image and text modalities. Therefore, we introduce the Cross-Modal Mask and Detail Alignment (CMDA) framework to address these challenges. Under this framework, our proposed Cross-Modal Mask Alignment module (CMA) semantically aligns features generated by supplementing randomly masked image/text with another modality text/image. Additionally, to narrow the gap between the image and text modalities, we designed the Cross-Modal Detail Alignment module (CDA), which establishes connections between images and texts and facilitates interactions between these two different modalities. Experimental results show that our model exhibits outstanding performance across multiple public datasets, i.e., CUHK-PEDES, ICFG-PEDES, and RSTPReID.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Modal Mask and Detail Alignment for Text-Based Person Retrieval

  • Ao Guo,
  • Xuan Liu,
  • Xianggan Liu,
  • Bingmeng Hu,
  • Jie Yuan,
  • Chiawei Chu

摘要

Text-based person retrieval primarily aims to retrieve the images of target persons represented by a given text query. In this task, how to effectively align images and text globally and locally is an important challenge. At the same time, since multiple modalities are involved, reducing the differences between different modalities is also an important challenge. Existing work has focused more on reconstructing an image rather than the semantic consistency between the image and text modalities. Therefore, we introduce the Cross-Modal Mask and Detail Alignment (CMDA) framework to address these challenges. Under this framework, our proposed Cross-Modal Mask Alignment module (CMA) semantically aligns features generated by supplementing randomly masked image/text with another modality text/image. Additionally, to narrow the gap between the image and text modalities, we designed the Cross-Modal Detail Alignment module (CDA), which establishes connections between images and texts and facilitates interactions between these two different modalities. Experimental results show that our model exhibits outstanding performance across multiple public datasets, i.e., CUHK-PEDES, ICFG-PEDES, and RSTPReID.