<p>In recent years, the crowd-counting system has strong strength to ensure safety in public places with the practical necessity of dense crowd analysis. However, until now, it has been difficult to obtain high-quality density maps due to complex background interference, congested crowd and large-scale variations. To address this issue, this paper presents a multi-scale feature fusion network (MSFFNet) based on CNN (convolution neural network), which is capable of detecting enough semantic features to understand crowds in sparse and highly congested scenes. In this method, a large majority of encoded features are fused adaptively rather than separated extraction components. Therefore, it enhances the ability to extend the range of receptive field sizes and reduce computation cost. MSSFNet is consists of three modules: grouped feature extractor, fusion block and decoder. The feature extractor is based on first 13 convolution layers of VGG16, which extract low-level features from crowd images. The fusion block computes the weights in each group from the contrast features and average them from convolutional layers later concatenated pooling layer with a feature map. The decoder capably extracts relevant information while enduring spatial resolution. Additionally, we designed two-stream module and semantic optimization module (SOM) with decoder which instantaneously enhance the crowd head positions and reduce background noises by re-weighting features. Extensive experiments on four public datasets (ShanghaiTech Part_A and Part_B, UCF_CC_50, UCF_QNRF and JHU-Crowd++), validate that MSFFNet can perform efficiently in complex background noises and capture head sizes in sparse, congested and various weather situation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MSFFNet: multi-scale feature fusion network with semantic optimization for crowd counting

  • Avinash Rohra,
  • Baoqun Yin,
  • Hazrat Bilal,
  • Aakash Kumar,
  • Munawar Ali,
  • Yang Li

摘要

In recent years, the crowd-counting system has strong strength to ensure safety in public places with the practical necessity of dense crowd analysis. However, until now, it has been difficult to obtain high-quality density maps due to complex background interference, congested crowd and large-scale variations. To address this issue, this paper presents a multi-scale feature fusion network (MSFFNet) based on CNN (convolution neural network), which is capable of detecting enough semantic features to understand crowds in sparse and highly congested scenes. In this method, a large majority of encoded features are fused adaptively rather than separated extraction components. Therefore, it enhances the ability to extend the range of receptive field sizes and reduce computation cost. MSSFNet is consists of three modules: grouped feature extractor, fusion block and decoder. The feature extractor is based on first 13 convolution layers of VGG16, which extract low-level features from crowd images. The fusion block computes the weights in each group from the contrast features and average them from convolutional layers later concatenated pooling layer with a feature map. The decoder capably extracts relevant information while enduring spatial resolution. Additionally, we designed two-stream module and semantic optimization module (SOM) with decoder which instantaneously enhance the crowd head positions and reduce background noises by re-weighting features. Extensive experiments on four public datasets (ShanghaiTech Part_A and Part_B, UCF_CC_50, UCF_QNRF and JHU-Crowd++), validate that MSFFNet can perform efficiently in complex background noises and capture head sizes in sparse, congested and various weather situation.