Temporal-Semantic Context Fusion for Robust Weakly Supervised Video Anomaly Detection
摘要
Video anomaly detection (VAD) has emerged as a vital and challenging task in computer vision, driven by the rapid advancement of surveillance videos. Weakly supervised learning is a critical branch in this field, with Multiple Instance Learning (MIL) standing out as the prevalent approach. In the weakly supervised VAD task, video context information provides crucial cues. However, the majority of studies in this field have primarily focused on the temporal context, neglecting the emphasis of semantic context in videos, which provides a deeper understanding of the observed video content. Therefore, this paper proposes a MIL-based framework for efficiently capturing video context information in weakly supervised VAD task. A GCN-based Temporal-Semantic context fusion module is employed to comprehensively capture both temporal and semantic context. Additionally, in order to enhance feature discrimination and learning robustness, our approach integrates highly effective anomaly classification learning and feature magnitude learning loss functions. Extensive experiments on two challenging datasets (i.e., ShanghaiTech and UCF-Crime) outperform some recent methods and demonstrate our approach’s effectiveness.