Visible-infrared person re-identification (VI-ReID) aims to associate pedestrian images captured by non-overlapping visible and infrared cameras. It faces two primary challenges related to modality discrepancy and intra-modality variations. To address the discrepancy between visible and infrared data, existing methods focus on generative-based methods that transform images from one modality to another or synthesize images in an intermediate modality. However, noise interference can compromise training stability and hinder effective modality alignment. In this paper, a novel Spectrum Adaptation Hierarchical Attention Network (SAHAN) is proposed to alleviate cross-modality discrepancies through color space transformation and enhance modality-shared representations with an attention-aggregation strategy. SAHAN comprises two key components: HSV Conversion Alignment (HCA) method and Hierarchical Attention Aggregation (HAA) strategy. Specifically, the HCA adjusts hue, saturation, and brightness in the HSV color space to reduce spectral distribution differences between visible and infrared images, thereby alleviating cross-modal discrepancies. Additionally, the HAA incorporates a multi-level attention aggregation mechanism to capture discriminative features at different stages. It then employs a multi-branch attention (MBA) module to further capture multi-scale information, including spatial and channel-wise features, thereby enhancing the capability to learn modality-shared representations. Extensive evaluations on the SYSU-MM01 and RegDB benchmarks demonstrate that our proposed method outperforms existing state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spectrum Adaptation Hierarchical Attention Network for Visible-Infrared Person Re-identification

  • Yong Wu,
  • Rongxi Zhou,
  • Hongchao Li,
  • Ze Zhou,
  • Guodui He,
  • Zhengyou Qin

摘要

Visible-infrared person re-identification (VI-ReID) aims to associate pedestrian images captured by non-overlapping visible and infrared cameras. It faces two primary challenges related to modality discrepancy and intra-modality variations. To address the discrepancy between visible and infrared data, existing methods focus on generative-based methods that transform images from one modality to another or synthesize images in an intermediate modality. However, noise interference can compromise training stability and hinder effective modality alignment. In this paper, a novel Spectrum Adaptation Hierarchical Attention Network (SAHAN) is proposed to alleviate cross-modality discrepancies through color space transformation and enhance modality-shared representations with an attention-aggregation strategy. SAHAN comprises two key components: HSV Conversion Alignment (HCA) method and Hierarchical Attention Aggregation (HAA) strategy. Specifically, the HCA adjusts hue, saturation, and brightness in the HSV color space to reduce spectral distribution differences between visible and infrared images, thereby alleviating cross-modal discrepancies. Additionally, the HAA incorporates a multi-level attention aggregation mechanism to capture discriminative features at different stages. It then employs a multi-branch attention (MBA) module to further capture multi-scale information, including spatial and channel-wise features, thereby enhancing the capability to learn modality-shared representations. Extensive evaluations on the SYSU-MM01 and RegDB benchmarks demonstrate that our proposed method outperforms existing state-of-the-art approaches.