Visible-infrared person re-identification (VI-ReID) aims to match the same pedestrian across modalities, which can make up for the limitations of single-modality person re-identification in real-world applications. Regarding the inherent modality discrepancy between the two modality images, existing research mainly focuses on generating target or intermediate modalities to transform cross-modality retrieval into single-modality retrieval, or designing network structures, loss functions, etc. to extract shared features of different modalities. However, generation-based methods easily introduce noise. Feature-based methods mostly use pooling to describe global features or horizontally divided partial features, which easily loses key information. To this end, we propose a Modality Mitigation and Diverse Part Awareness ( \(M^2 DPA\) ) network framework for VI-ReID. This framework introduces the modality mitigation module and the diverse part awareness module to mitigate modality discrepancy and extract local features containing more key information. Specifically, facing the modality discrepancy problem, we do not focus on generation, but recognize the effectiveness of Instance Normalization and utilize channel attention to guide Instance Normalization to eliminate modality information in feature maps while retaining discriminability. Then, in order to extract fine-grained part features, we design a diverse part awareness module. It exploits the correlation between pixel features to aggregate partial features with learnable part prototypes as reference. This pixel-level approach minimizes the loss of critical information. Comprehensive experimental results on SYSU-MM01 and RegDB datasets show that our \(M^2 DPA\) has good performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modality Mitigation And Diverse Part Awareness for Visible-Infrared Person Re-identification

  • Meiling Zhang,
  • Xin Li,
  • Qiang Wang,
  • Hubo Guo,
  • Zhihong Huang

摘要

Visible-infrared person re-identification (VI-ReID) aims to match the same pedestrian across modalities, which can make up for the limitations of single-modality person re-identification in real-world applications. Regarding the inherent modality discrepancy between the two modality images, existing research mainly focuses on generating target or intermediate modalities to transform cross-modality retrieval into single-modality retrieval, or designing network structures, loss functions, etc. to extract shared features of different modalities. However, generation-based methods easily introduce noise. Feature-based methods mostly use pooling to describe global features or horizontally divided partial features, which easily loses key information. To this end, we propose a Modality Mitigation and Diverse Part Awareness ( \(M^2 DPA\) ) network framework for VI-ReID. This framework introduces the modality mitigation module and the diverse part awareness module to mitigate modality discrepancy and extract local features containing more key information. Specifically, facing the modality discrepancy problem, we do not focus on generation, but recognize the effectiveness of Instance Normalization and utilize channel attention to guide Instance Normalization to eliminate modality information in feature maps while retaining discriminability. Then, in order to extract fine-grained part features, we design a diverse part awareness module. It exploits the correlation between pixel features to aggregate partial features with learnable part prototypes as reference. This pixel-level approach minimizes the loss of critical information. Comprehensive experimental results on SYSU-MM01 and RegDB datasets show that our \(M^2 DPA\) has good performance.