Video anomaly detection(VAD) aims to learn the normal appearance and motion patterns of video data and identify abnormal behaviours that deviate from these expected patterns. Existing methods employ easily fusion strategy to combine appearance and motion features. However, these methods overlook the characteristics of motion features in VAD, resulting in ineffective associations between appearance and motion features. In addition, most approaches employ coarse-grained modeling, which is inadequate for capturing intrinsic representations of normal behaviors. Therefore, we propose a multi-scale differential perception network(MSDPN). Firstly, we propose the differential perception fusion block(DPFB), which employs differential perception attention compute motion-salient regions through differential comparison and attention, enhancing the model’s sensitivity to critical motion cues. Secondly, we design a progressive multi-scale differential perception fusion strategy to establish connections between appearance and motion features across multiple scales, enriching the motion representation of the appearance branch. Finally, we design a memory network with feature refined fusion mechanism(MFRM), which removes redundant information through refined processing of deep features, enabling memory network to get more discriminative prototypes. Extensive experiments demonstrate that our MSDPN outperforms the state-of-the-art methods, achieving AUC improvements of at least 0.2%, 1.6%, and 1.0% on the Ped2, Avenue, and ShanghaiTech datasets, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale Differential Perception Network for Video Anomaly Detection

  • Min Jiang,
  • Weilong Wang,
  • Jun Kong

摘要

Video anomaly detection(VAD) aims to learn the normal appearance and motion patterns of video data and identify abnormal behaviours that deviate from these expected patterns. Existing methods employ easily fusion strategy to combine appearance and motion features. However, these methods overlook the characteristics of motion features in VAD, resulting in ineffective associations between appearance and motion features. In addition, most approaches employ coarse-grained modeling, which is inadequate for capturing intrinsic representations of normal behaviors. Therefore, we propose a multi-scale differential perception network(MSDPN). Firstly, we propose the differential perception fusion block(DPFB), which employs differential perception attention compute motion-salient regions through differential comparison and attention, enhancing the model’s sensitivity to critical motion cues. Secondly, we design a progressive multi-scale differential perception fusion strategy to establish connections between appearance and motion features across multiple scales, enriching the motion representation of the appearance branch. Finally, we design a memory network with feature refined fusion mechanism(MFRM), which removes redundant information through refined processing of deep features, enabling memory network to get more discriminative prototypes. Extensive experiments demonstrate that our MSDPN outperforms the state-of-the-art methods, achieving AUC improvements of at least 0.2%, 1.6%, and 1.0% on the Ped2, Avenue, and ShanghaiTech datasets, respectively.