Enhancing cross-modal person reidentification: multi-scale feature alignment and optimization
摘要
Person reidentification across visible and infrared spectra is a critical task in surveillance and security, facing challenges in bridging the modality gap. This study introduces a multi-scale perception and semantic understanding (MPSU) framework to enhance cross-modal feature consistency. The MPSU framework incorporates a lightweight multi-scale (LMS) module for rich feature representation and a center sample similarity aggregation (CSSA) module for capturing cross-modal correlations. Experimental results on SYSU-MM01 and RegDB datasets demonstrate significant performance improvements, achieving a Rank-1 accuracy of 78.6% and mAP of 76.1% on SYSU-MM01, and 94.8% Rank-1 accuracy and 91.5% mAP on RegDB, showcasing superior cross-modal matching capabilities.