<p>The inherent complexity of visible-infrared person re-identification is characterized by significant intra-class variance and pronounced inter-modal disparities. Existing approaches address these challenges by constructing comprehensive data representations through joint learning of multi-modal samples or cross-modal transformation techniques. However, the lack of a dynamic modulation mechanism limits their ability to adapt modality-specific features effectively, thereby constraining the generalizability of the shared feature space. This limitation results in a shared space that lacks the robustness necessary for effective generalization. To address these issues, we introduce the dual dynamic modality alignment network, a novel framework designed to dynamically calibrate the significance of modality-specific features, optimizing critical data extraction while minimizing reliance on extraneous information. Central to our approach is the class-aware modality hybrid-assisted generator, which conceptualizes the multimodal contrastive representation space as nodes, integrates diverse contrastive representations, and interlinks isolated representations to explore a wider array of contrastive relationships between modalities. Additionally, we propose an auxiliary modal identity center alignment loss that refines feature distribution and reduces divergence between visible and infrared image representations. Extensive evaluation on the SYSU-MM01 and RegDB datasets demonstrates the superior performance of our method, emphasizing its efficacy in creating a more discriminative and balanced shared feature space.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modality-agnostic learning for robust visible-infrared person re-identification

  • Shengrong Gong,
  • Shuomin Li,
  • Gengsheng Xie,
  • Yufeng Yao,
  • Shan Zhong

摘要

The inherent complexity of visible-infrared person re-identification is characterized by significant intra-class variance and pronounced inter-modal disparities. Existing approaches address these challenges by constructing comprehensive data representations through joint learning of multi-modal samples or cross-modal transformation techniques. However, the lack of a dynamic modulation mechanism limits their ability to adapt modality-specific features effectively, thereby constraining the generalizability of the shared feature space. This limitation results in a shared space that lacks the robustness necessary for effective generalization. To address these issues, we introduce the dual dynamic modality alignment network, a novel framework designed to dynamically calibrate the significance of modality-specific features, optimizing critical data extraction while minimizing reliance on extraneous information. Central to our approach is the class-aware modality hybrid-assisted generator, which conceptualizes the multimodal contrastive representation space as nodes, integrates diverse contrastive representations, and interlinks isolated representations to explore a wider array of contrastive relationships between modalities. Additionally, we propose an auxiliary modal identity center alignment loss that refines feature distribution and reduces divergence between visible and infrared image representations. Extensive evaluation on the SYSU-MM01 and RegDB datasets demonstrates the superior performance of our method, emphasizing its efficacy in creating a more discriminative and balanced shared feature space.