<p>Person re-identification is a challenging problem in video surveillance because of the variations in people poses across camera views, illumination changes, and occlusions. To address these challenges, we propose a novel framework that strategically integrates a pose-aware attention-enhanced generative adversarial network (AE-GAN) with an attention-driven feature fusion strategy for re-identification. The proposed AE-GAN synthesizes camera-specific pose-transformed images of people appearing in the next camera, thus bridging the viewpoint gap and minimizing pose-related discrepancies. This generative augmentation also reduces the dependency on extensive cross-view annotations. We process the synthesized and real images using a ResNet-50 backbone enhanced with multi-head self-attention modules that highlight discriminative features, such as clothing details, while suppressing irrelevant backgrounds. A subsequent fusion layer aggregates multi-scale features. We optimize the system using a combined classification and contrastive loss to improve inter-identity separability. Experimental results on publicly available challenging MARS and CUHK03-NP datasets demonstrate that our method outperforms the state-of-the-art ReID approaches and offers a scalable and robust solution for complex surveillance environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

View-invariant person re-identification in camera network by pose-aware AE-GAN and attention-driven feature fusion

  • Syed Fahad Tahir,
  • Muhammad Usman,
  • Anas Mustafa,
  • Muhammad Shaheer Ali,
  • Labiba Gillani Fahad

摘要

Person re-identification is a challenging problem in video surveillance because of the variations in people poses across camera views, illumination changes, and occlusions. To address these challenges, we propose a novel framework that strategically integrates a pose-aware attention-enhanced generative adversarial network (AE-GAN) with an attention-driven feature fusion strategy for re-identification. The proposed AE-GAN synthesizes camera-specific pose-transformed images of people appearing in the next camera, thus bridging the viewpoint gap and minimizing pose-related discrepancies. This generative augmentation also reduces the dependency on extensive cross-view annotations. We process the synthesized and real images using a ResNet-50 backbone enhanced with multi-head self-attention modules that highlight discriminative features, such as clothing details, while suppressing irrelevant backgrounds. A subsequent fusion layer aggregates multi-scale features. We optimize the system using a combined classification and contrastive loss to improve inter-identity separability. Experimental results on publicly available challenging MARS and CUHK03-NP datasets demonstrate that our method outperforms the state-of-the-art ReID approaches and offers a scalable and robust solution for complex surveillance environments.