错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

cMDTPS: Comprehensive Masked Modality Modeling with Improved Similarity Distribution Matching Loss for Text-based Person Search

  • Anh D. Nguyen,
  • Dang H. Pham,
  • Duc M. Nguyen,
  • Hoa N. Nguyen

摘要

The goal of text-based person search is to use a textual description query to find the required individual within an image gallery. There are two major obstacles to this task: (i) how to bridge the gap between the textual and visual modalities in the feature space and (ii) how to improve the image representation by removing the focus of the model on unnecessary regions from the image as the same way we do on text input. To address these challenges, we propose cMDTPS, a comprehensive masked modality modeling method with improved similarity distribution matching loss. The proposed method consists of two components: (i) an improved cross-modality alignment loss that minimizes the distance between the distributions of the same person in different modalities, and (ii) a comprehensive masked modality modeling method that helps the model focus on important parts of inputs. We evaluate our proposed method on three popular benchmarks: CUHK-PEDES, ICFG-PEDES, and RSTPReid, and show that it outperforms the state-of-the-art methods.