Weakly Supervised Video Anomaly Detection (WS-VAD) aims to utilize Multiple Instance Learning (MIL) to assign anomaly scores to snippets based on pretrained video features. However, the feature extractors pretrained on other domain may produce suboptimal representations and visual-semantic misalignment in the target domain, resulting in inaccurate anomaly rankings within the MIL framework. To address these challenges, we propose a novel Semantic-Enhanced WS-VAD framework(SE-WSVAD), which leverages high-level semantic descriptions generated by Visual Language Models (VLMs). Within the proposed framework, we introduce three key modules: the Semantic Noise Filter (SNF) to filter redundant and noisy semantic information, the Temporal Context Summarizer (TCS) to refine snippet-level semantic representations, and the Semantic-Aware Video Alignment Module (SAVAM) to align visual features with semantic context via a cross-attention mechanism. Extensive experiments on ShanghaiTech and UBnormal datasets demonstrate the efficacy of SE-WSVAD, achieving significant improvements in anomaly detection performance over state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Filter, Summarize, Align: Learning the Semantic Guided Weakly Supervised Video Anomaly Detection

  • Hengpeng Xu,
  • Hua Wang,
  • Muye Yue,
  • Zhengtao Li

摘要

Weakly Supervised Video Anomaly Detection (WS-VAD) aims to utilize Multiple Instance Learning (MIL) to assign anomaly scores to snippets based on pretrained video features. However, the feature extractors pretrained on other domain may produce suboptimal representations and visual-semantic misalignment in the target domain, resulting in inaccurate anomaly rankings within the MIL framework. To address these challenges, we propose a novel Semantic-Enhanced WS-VAD framework(SE-WSVAD), which leverages high-level semantic descriptions generated by Visual Language Models (VLMs). Within the proposed framework, we introduce three key modules: the Semantic Noise Filter (SNF) to filter redundant and noisy semantic information, the Temporal Context Summarizer (TCS) to refine snippet-level semantic representations, and the Semantic-Aware Video Alignment Module (SAVAM) to align visual features with semantic context via a cross-attention mechanism. Extensive experiments on ShanghaiTech and UBnormal datasets demonstrate the efficacy of SE-WSVAD, achieving significant improvements in anomaly detection performance over state-of-the-art methods.