Video anomaly detection distinguishes between normal and abnormal events based on differences in feature representations. However, most existing methods are still limited in their video representation capabilities, and there is an imbalance in the data volumes of normal and abnormal events. To address these issues, we propose a multi-instance weakly supervised training framework that efficiently optimizes task-specific discriminative representations using video-level label information. The core components of the framework include: (1) an axial attention-enhanced feature extractor, which integrates global video information and maintains long-range spatial dependencies; (2) a fused BLSTM module, which captures spatiotemporal features before and after abnormal events, automatically focusing on anomalous regions within frames and their contextual information; (3) hard negative feature set, which selects the highest-scoring segments from normal video clips, dynamically updated using a cross-period sampling strategy to ensure the classifier is continuously optimized with the latest challenging samples. Experimental results show that our method outperforms many existing approaches in terms of accuracy and false positive rate, achieving a frame-level accuracy of 94.32% on the ShanghaiTech dataset and a false positive rate as low as 0.02% on the UCF-Crime dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multiple Instance Weakly-Supervised Training Framework for Video Anomaly Detection

  • Ying Hu,
  • Yuanyao Lu

摘要

Video anomaly detection distinguishes between normal and abnormal events based on differences in feature representations. However, most existing methods are still limited in their video representation capabilities, and there is an imbalance in the data volumes of normal and abnormal events. To address these issues, we propose a multi-instance weakly supervised training framework that efficiently optimizes task-specific discriminative representations using video-level label information. The core components of the framework include: (1) an axial attention-enhanced feature extractor, which integrates global video information and maintains long-range spatial dependencies; (2) a fused BLSTM module, which captures spatiotemporal features before and after abnormal events, automatically focusing on anomalous regions within frames and their contextual information; (3) hard negative feature set, which selects the highest-scoring segments from normal video clips, dynamically updated using a cross-period sampling strategy to ensure the classifier is continuously optimized with the latest challenging samples. Experimental results show that our method outperforms many existing approaches in terms of accuracy and false positive rate, achieving a frame-level accuracy of 94.32% on the ShanghaiTech dataset and a false positive rate as low as 0.02% on the UCF-Crime dataset.