<p>Visible-Infrared Person Re-Identification (VI Re-ID) is designed to match visible pedestrian images with infrared pedestrian images. The modality discrepancy between visible and infrared images is the biggest challenge for VI Re-ID. Existing VI Re-ID approaches mostly weaken the effect of modality discrepancy by extracting discriminative features, while ignoring the ability of models to adapt to modality variations. To address this issue, we construct a CNN-Transformer hybrid model (CTHM) for VI Re-ID, which mainly consists of a dual-stream ResNet-50 network, a dual-stream VIT network, and a multi-modal feature fusion module (MFFM). In particular, we design a multi-modal multi-stage training strategy (MMTS). Given visible images and infrared images, MMTS first uses Cycle GAN to generate fake infrared images and fake visible images. Then, MMTS sequentially employs visible images and fake visible images, infrared images and fake infrared images, fake visible images and fake infrared images, and visible images and infrared images, to train CTHM in stages in order to progressively improve its adaptive ability to modality variations, thus ensuring that CTHM can obtain more discriminative modality-invariant features and better attenuates the influence of modality discrepancy. We conduct extensive experiments on two benchmark datasets, SYSU-MM01 and RegDB, and the results show that our method reaches the current advanced level.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A CNN-transformer hybrid model and a multi-modal multi-stage training strategy for visible-infrared person re-identification

  • Xinxin Hao,
  • Haishun Du,
  • Jiangtao Guo,
  • Jieru Li

摘要

Visible-Infrared Person Re-Identification (VI Re-ID) is designed to match visible pedestrian images with infrared pedestrian images. The modality discrepancy between visible and infrared images is the biggest challenge for VI Re-ID. Existing VI Re-ID approaches mostly weaken the effect of modality discrepancy by extracting discriminative features, while ignoring the ability of models to adapt to modality variations. To address this issue, we construct a CNN-Transformer hybrid model (CTHM) for VI Re-ID, which mainly consists of a dual-stream ResNet-50 network, a dual-stream VIT network, and a multi-modal feature fusion module (MFFM). In particular, we design a multi-modal multi-stage training strategy (MMTS). Given visible images and infrared images, MMTS first uses Cycle GAN to generate fake infrared images and fake visible images. Then, MMTS sequentially employs visible images and fake visible images, infrared images and fake infrared images, fake visible images and fake infrared images, and visible images and infrared images, to train CTHM in stages in order to progressively improve its adaptive ability to modality variations, thus ensuring that CTHM can obtain more discriminative modality-invariant features and better attenuates the influence of modality discrepancy. We conduct extensive experiments on two benchmark datasets, SYSU-MM01 and RegDB, and the results show that our method reaches the current advanced level.