View-invariant person re-identification in camera network by pose-aware AE-GAN and attention-driven feature fusion
摘要
Person re-identification is a challenging problem in video surveillance because of the variations in people poses across camera views, illumination changes, and occlusions. To address these challenges, we propose a novel framework that strategically integrates a pose-aware attention-enhanced generative adversarial network (AE-GAN) with an attention-driven feature fusion strategy for re-identification. The proposed AE-GAN synthesizes camera-specific pose-transformed images of people appearing in the next camera, thus bridging the viewpoint gap and minimizing pose-related discrepancies. This generative augmentation also reduces the dependency on extensive cross-view annotations. We process the synthesized and real images using a ResNet-50 backbone enhanced with multi-head self-attention modules that highlight discriminative features, such as clothing details, while suppressing irrelevant backgrounds. A subsequent fusion layer aggregates multi-scale features. We optimize the system using a combined classification and contrastive loss to improve inter-identity separability. Experimental results on publicly available challenging MARS and CUHK03-NP datasets demonstrate that our method outperforms the state-of-the-art ReID approaches and offers a scalable and robust solution for complex surveillance environments.