Multi-granularity and Multi-modal Prompt Learning for Person Re-Identification
摘要
Pre-trained vision-language models, such as CLIP, are driving advancements in person re-identification by mining semantic information. Current approaches utilize globally learnable textual prompts to generate coarse-level, holistic yet ambiguous descriptions of individuals, which are then served as constraints for the fine-tuning of CLIP to learn visual representations of the pedestrians. However, relying exclusively on this type of prompt learning for the text encoder of CLIP overlooks the crucial fine-grained details of individuals and fails to model the necessary downstream adaptation capacity for the image encoder of CLIP. To address these limitations, we propose a novel Multi-Granularity and Multi-Modal Prompt Learning (MMPL) for person re-identification to fully unleash the substantial potential inherent in CLIP for acquiring discriminative representations. The MMPL encompasses a two-stage training procedure. In the first training stage, MMPL meticulously orchestrates the Hierarchical Prompt Learning (HPL) to refine crucial and distinctive information from hierarchical patch-level visual features. Aligning textual prompts with these subtle visual cues across diverse granularities, this process establishes patch-to-token level correspondences, ultimately yielding the creation of high-fidelity multi-granularity textual prompts. In the second training stage, MMPL integrates Collaborative Prompt Learning (CPL), generating supplementary visual prompts and fostering multi-modal interactive learning to aid CLIP’s image encoder in narrowing the semantic gap between modalities, leveraging CLIP’s extensive multi-modal knowledge to enhance feature representation. Comprehensive experimental evaluation across four widely recognized person re-identification benchmarks substantiates the effectiveness of our MMPL.