Continual adaptation Person re-identification via vision-language fusion with enhanced annotation robustness
摘要
Person re-identification (ReID) in real-world multimodal settings can benefit substantially from the integration of visual and textual modalities. However, this task remains challenging due to modality misalignment, noisy textual annotations, and performance degradation under domain shifts during deployment. To address these challenges, we propose EAR-ReID (Enhanced Annotation Robustness for Re-ID), a robust and adaptive cross-modal ReID framework that introduces three key innovations. First, we design a Dual-Stream Cross-Attention Fusion (DCAF) module to enable fine-grained semantic alignment between image and caption features via bidirectional cross-attention, enhancing identity-discriminative representation learning. Second, we present a Vision-Guided Noisy Text Suppression (VNTS) mechanism, which estimates token reliability using image-guided attention and employs a RINCE loss to suppress the influence of hallucinated or misleading captions. Third, to support continual adaptation without catastrophic forgetting, we introduce a Continual Alignment via Elastic Weight Consolidation (CA-EWC) strategy, which preserves cross-modal knowledge while incrementally updating the model to accommodate new data distributions. Extensive experiments on several benchmark datasets demonstrate that EAR-ReID achieves state-of-the-art performance and maintains strong robustness under both static and continually evolving environments.