The real-world applicability of automated violence recognition systems has drawn much attention from researchers. The current techniques for recognizing violence are centered on creating efficient models that can predict violent events quickly and accurately in real-time. However, early violence prediction, which is crucial for real-time systems, is not considered in these methods. In this paper, we present an early violence prediction method that can accurately predict violent activities from partially observed video frames. We propose a two-stream architecture which employs our proposed efficient Squeeze-Excitation ShuffleNet (SESNet) model that effectively extracts spatial and temporal features. We leverage spatio-temporal and channel-wise squeeze and excitation to incorporate attention information into the ShuffleNet V2 architecture. To enable early violence recognition, we train our model in a teacher-student framework, where the teacher model trained on full-length videos distils privileged information to the student model, which has access to partial videos. For this purpose, we introduce a novel multi-teacher importance preservation learning methodology which can effectively distill important features from multiple teacher networks. We evaluate our approach on the challenging RWF-2000 public violence recognition dataset. Experimental results show that our teacher-student training framework performed well for early violence prediction. Additionally, our proposed model also outperforms several state-of-the-art violence recognition methods on full-length videos.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-teacher Importance Preserving Knowledge Distillation for Early Violence Prediction

  • Suvramalya Basak,
  • Aditya Vaishy,
  • Anjali Gautam

摘要

The real-world applicability of automated violence recognition systems has drawn much attention from researchers. The current techniques for recognizing violence are centered on creating efficient models that can predict violent events quickly and accurately in real-time. However, early violence prediction, which is crucial for real-time systems, is not considered in these methods. In this paper, we present an early violence prediction method that can accurately predict violent activities from partially observed video frames. We propose a two-stream architecture which employs our proposed efficient Squeeze-Excitation ShuffleNet (SESNet) model that effectively extracts spatial and temporal features. We leverage spatio-temporal and channel-wise squeeze and excitation to incorporate attention information into the ShuffleNet V2 architecture. To enable early violence recognition, we train our model in a teacher-student framework, where the teacher model trained on full-length videos distils privileged information to the student model, which has access to partial videos. For this purpose, we introduce a novel multi-teacher importance preservation learning methodology which can effectively distill important features from multiple teacher networks. We evaluate our approach on the challenging RWF-2000 public violence recognition dataset. Experimental results show that our teacher-student training framework performed well for early violence prediction. Additionally, our proposed model also outperforms several state-of-the-art violence recognition methods on full-length videos.