Filter, Summarize, Align: Learning the Semantic Guided Weakly Supervised Video Anomaly Detection
摘要
Weakly Supervised Video Anomaly Detection (WS-VAD) aims to utilize Multiple Instance Learning (MIL) to assign anomaly scores to snippets based on pretrained video features. However, the feature extractors pretrained on other domain may produce suboptimal representations and visual-semantic misalignment in the target domain, resulting in inaccurate anomaly rankings within the MIL framework. To address these challenges, we propose a novel Semantic-Enhanced WS-VAD framework(SE-WSVAD), which leverages high-level semantic descriptions generated by Visual Language Models (VLMs). Within the proposed framework, we introduce three key modules: the Semantic Noise Filter (SNF) to filter redundant and noisy semantic information, the Temporal Context Summarizer (TCS) to refine snippet-level semantic representations, and the Semantic-Aware Video Alignment Module (SAVAM) to align visual features with semantic context via a cross-attention mechanism. Extensive experiments on ShanghaiTech and UBnormal datasets demonstrate the efficacy of SE-WSVAD, achieving significant improvements in anomaly detection performance over state-of-the-art methods.