Cross-modality person re-identification, particularly between daytime RGB images and nighttime infrared (IR) images, is one of the challenges in image retrieval tasks on the visual database. Cross-modality feature alignment is one typical solution in current works. However, the significant difference between RGB and IR modalities makes it difficult to align the two features directly. To reduce the modality gap, we introduce an intermediate modality, a grayscale image, which can be generated from RGB with the same label. The grayscale image retains the structural information of the RGB modality and exhibits a visual style similar to the IR image. To improve the performance of Re-ID in visible-infrared cross-modality tasks, we propose a Tri-modality Collaborative Learning model (TCL). There are two modules in TCL, the tri-modality joint feature extraction module and the modality-specific mean teaching classifier module. In the feature extraction module, multiple channels with different modalities are trained to learn modality-independent high-dimensional shared features. The classifier module is designed to ignore modality-specific information. We evaluate the effectiveness of TCL in the public dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tri-modality Collaborative Learning for Person Re-identification

  • Shizhuo Deng,
  • Qingyuan Yang,
  • Zhibin Yang,
  • Dongyue Chen,
  • Yu Yang,
  • Hao Wang

摘要

Cross-modality person re-identification, particularly between daytime RGB images and nighttime infrared (IR) images, is one of the challenges in image retrieval tasks on the visual database. Cross-modality feature alignment is one typical solution in current works. However, the significant difference between RGB and IR modalities makes it difficult to align the two features directly. To reduce the modality gap, we introduce an intermediate modality, a grayscale image, which can be generated from RGB with the same label. The grayscale image retains the structural information of the RGB modality and exhibits a visual style similar to the IR image. To improve the performance of Re-ID in visible-infrared cross-modality tasks, we propose a Tri-modality Collaborative Learning model (TCL). There are two modules in TCL, the tri-modality joint feature extraction module and the modality-specific mean teaching classifier module. In the feature extraction module, multiple channels with different modalities are trained to learn modality-independent high-dimensional shared features. The classifier module is designed to ignore modality-specific information. We evaluate the effectiveness of TCL in the public dataset.