Tri-modality Collaborative Learning for Person Re-identification
摘要
Cross-modality person re-identification, particularly between daytime RGB images and nighttime infrared (IR) images, is one of the challenges in image retrieval tasks on the visual database. Cross-modality feature alignment is one typical solution in current works. However, the significant difference between RGB and IR modalities makes it difficult to align the two features directly. To reduce the modality gap, we introduce an intermediate modality, a grayscale image, which can be generated from RGB with the same label. The grayscale image retains the structural information of the RGB modality and exhibits a visual style similar to the IR image. To improve the performance of Re-ID in visible-infrared cross-modality tasks, we propose a Tri-modality Collaborative Learning model (TCL). There are two modules in TCL, the tri-modality joint feature extraction module and the modality-specific mean teaching classifier module. In the feature extraction module, multiple channels with different modalities are trained to learn modality-independent high-dimensional shared features. The classifier module is designed to ignore modality-specific information. We evaluate the effectiveness of TCL in the public dataset.