Weakly supervised video anomaly detection (WVAD) aims to detect where abnormal events occur in videos using only video-level labels during training. Many existing WVAD methods concentrate on learning comprehensive representations for each frame, making them susceptible to interference from irrelevant backgrounds or scenes. In this paper, we address this challenge by introducing disentangled representation learning to WVAD, presenting a new modular component named the Event-Centric Disentangler (ECD). Specifically, our ECD incorporates an Event Focus Attention (EFA) module, estimating channel-wise attention to focus on event representations while filtering out irrelevant information from backgrounds or scenes. Alongside EFA, a Frame Importance Allocator (FIA) learns frame-wise weighting factors to aggregate frame-level predictions for the generation of video-level predictions. Furthermore, we introduce a video-level contrastive loss to provide disentanglement supervision for ECD training, within the constraints of weakly supervised settings. Integrating ECD into existing WVAD methods, we achieve state-of-the-art performance on two benchmark datasets, with 87.53% AUC and 81.61% AP on the UCF-Crime and XD-Violence datasets, respectively. Our code is available at https://github.com/quyiii/ECD .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ECD: Event-Centric Disentangler for Weakly Supervised Video Anomaly Detection

  • Yi Qu,
  • Yixuan Zhou,
  • Guofeng Yi

摘要

Weakly supervised video anomaly detection (WVAD) aims to detect where abnormal events occur in videos using only video-level labels during training. Many existing WVAD methods concentrate on learning comprehensive representations for each frame, making them susceptible to interference from irrelevant backgrounds or scenes. In this paper, we address this challenge by introducing disentangled representation learning to WVAD, presenting a new modular component named the Event-Centric Disentangler (ECD). Specifically, our ECD incorporates an Event Focus Attention (EFA) module, estimating channel-wise attention to focus on event representations while filtering out irrelevant information from backgrounds or scenes. Alongside EFA, a Frame Importance Allocator (FIA) learns frame-wise weighting factors to aggregate frame-level predictions for the generation of video-level predictions. Furthermore, we introduce a video-level contrastive loss to provide disentanglement supervision for ECD training, within the constraints of weakly supervised settings. Integrating ECD into existing WVAD methods, we achieve state-of-the-art performance on two benchmark datasets, with 87.53% AUC and 81.61% AP on the UCF-Crime and XD-Violence datasets, respectively. Our code is available at https://github.com/quyiii/ECD .