Visible-infrared person re-identification (VI-ReID) has captured growing attention for its applications in surveillance within low-light environments. Due to the substantial modality discrepancy and pedestrian variations, VI-ReID remains a challenging task. In this paper, a weighted feature supplementation and feature alignment network (WFSFA-Net) is presented to tackle the primary challenges in VI-ReID. The proposed approach consists of two modules - the Weighted Feature Supplementation (WFS) module and the Cross-modal Feature Alignment (CMFA) module. WFS module can generate supplementary embeddings to mine informative representations to narrow the modality gap. CMFA module mines structural relationships between multi-modal features of the same pedestrian and then aligns these features of the two modalities by using the shortest path algorithm. This process can enhance the model’s robustness and generalization against pedestrian variations. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our approach, outperforming state-of-the-art methods by more than 3% in accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

WFSFA-Net: Weighted Feature Supplementation and Cross-Modal Feature Alignment for Visible-Infrared Person Re-identification

  • Huangyao Deng,
  • Zili Zhang,
  • Wenxin Dong,
  • Chi Zhang

摘要

Visible-infrared person re-identification (VI-ReID) has captured growing attention for its applications in surveillance within low-light environments. Due to the substantial modality discrepancy and pedestrian variations, VI-ReID remains a challenging task. In this paper, a weighted feature supplementation and feature alignment network (WFSFA-Net) is presented to tackle the primary challenges in VI-ReID. The proposed approach consists of two modules - the Weighted Feature Supplementation (WFS) module and the Cross-modal Feature Alignment (CMFA) module. WFS module can generate supplementary embeddings to mine informative representations to narrow the modality gap. CMFA module mines structural relationships between multi-modal features of the same pedestrian and then aligns these features of the two modalities by using the shortest path algorithm. This process can enhance the model’s robustness and generalization against pedestrian variations. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our approach, outperforming state-of-the-art methods by more than 3% in accuracy.