Self-attention Network with Dynamic Perception for Crowd Counting
摘要
Current crowd-counting networks face large-scale crowd scenes with changing distributions. To overcome this challenge, we propose an end-to-end crowd-counting network which uses fine-grained optimization and dynamic perception to accurate density map. The proposed network consists of VGG19 for feature encoding, a dual fine-grained attention module (DFAM) and an enhancement dynamic visual field attention module (EDVFM). DFAM optimizes feature maps at the pixel level in both spatial and channel dimensions, thereby reducing the impact of environmental changes and perspective distortion. EDVFM enables the network to adjust its feature extraction strategy dynamically, based on actual crowd distribution, improving its detection capability for randomly changing crowd scenes. Experimental results on four publicly available datasets (ShanghaiTech, UCF-CC-50, World-Expo ‘10, and Mall) show that the proposed network outperforms other networks.