Weakly supervised video anomaly detection plays a pivotal role in widely deployed surveillance systems. Most existing methods are based on the multi-instance learning paradigm, determining the predicted label of a video based on segments with higher prediction scores. During training, the model predominantly focuses on the segment with the highest anomaly score or the top-k highest-scored segments, neglecting the other segments. This bias towards certain segment features during training results in missed and false detections of anomaly segments, subsequently impacting the performance of video anomaly detection. In this paper, we introduce a contrastive loss strategy to uncover easily overlooked normal and abnormal segments, enhancing the distinction between normal and abnormal segments through contrastive loss. Additionally, we propose a multi-scale feature fusion approach to learn features from different scales of videos and integrate them into a more comprehensive feature representation to accommodate the diversity of anomaly events. Experimental results on the UCF-Crime and XD-Violence datasets validate the efficacy of our proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Weakly Supervised Video Anomaly Detection Method Based on Multi-scale Feature Fusion and Contrastive Loss

  • Kun Yang,
  • Zhiming Luo,
  • Shaozi Li

摘要

Weakly supervised video anomaly detection plays a pivotal role in widely deployed surveillance systems. Most existing methods are based on the multi-instance learning paradigm, determining the predicted label of a video based on segments with higher prediction scores. During training, the model predominantly focuses on the segment with the highest anomaly score or the top-k highest-scored segments, neglecting the other segments. This bias towards certain segment features during training results in missed and false detections of anomaly segments, subsequently impacting the performance of video anomaly detection. In this paper, we introduce a contrastive loss strategy to uncover easily overlooked normal and abnormal segments, enhancing the distinction between normal and abnormal segments through contrastive loss. Additionally, we propose a multi-scale feature fusion approach to learn features from different scales of videos and integrate them into a more comprehensive feature representation to accommodate the diversity of anomaly events. Experimental results on the UCF-Crime and XD-Violence datasets validate the efficacy of our proposed method.