GTF-SAM: Graph-temporal fusion and scene-aware margin for weakly supervised video anomaly detection
摘要
Weakly Supervised Video Anomaly Detection (WSVAD) aims to identify anomalous events in videos using only video-level labels, providing an efficient solution for automated surveillance. However, existing WSVAD methods face two key challenges. First, anomalies are context-dependent; for instance, running is normal in a park but anomalous in a library, complicating detection across diverse scenarios. Second, existing methods often employ attention mechanisms in high-dimensional feature spaces, leading to increased parameters and computational overhead, which hinders practical deployment. To address these challenges, we propose a novel WSVAD framework integrating the Graph-Enhanced Temporal Fusion (GTF) module and the Scene-Aware Margin (SAM) module. The GTF module operates in a low-dimensional feature space, combining attention and graph-based propagation to efficiently model temporal dependencies with reduced parameters. The SAM module enhances detection by enforcing scene-specific feature separation, distinguishing normal and anomalous events in varied contexts. Extensive experiments on UCF-Crime and XD-Violence datasets demonstrate that our framework achieves superior performance and efficiency, validated by comprehensive ablation studies.