A Multi-scale Feature Extraction and Alignment Method for Cross-Modal Person Re-Identification
摘要
In the field of computer vision, cross-modal person re-identification has garnered significant attention due to its applications in nighttime surveillance and multimodal security systems. However, achieving high recognition accuracy remains a challenge due to inherent differences in physical properties, environmental conditions across modalities, and difficulties in effectively utilizing multi-scale information. To address these issues, this paper proposes a novel deep learning architecture. The framework leverages a multi-scale enhancement module to capture features at various scales, a deep feature synthesis module to integrate hierarchical information across layers, and a modal feature harmonization module to eliminate modality-specific characteristics. By aligning features effectively between modalities, the proposed method enhances feature compatibility and adaptability under diverse conditions. Experimental results on three public datasets demonstrate that the proposed approach achieves superior rank-1 recognition accuracy compared to state-of-the-art methods. Ablation studies further validate the individual contributions of each module, showcasing the effectiveness and robustness of the proposed architecture in addressing cross-modal challenges.