<p>Weakly-Supervised Video Anomaly Detection (WSVAD) is a critical technology in intelligent surveillance, enabling precise temporal localization of anomalies using only video-level labels. While CLIP-based approaches have significantly advanced the field, two key limitations persist: (1) inadequate temporal modeling capabilities, and (2) suboptimal utilization of textual information. These constraints stem fundamentally from CLIP’s static image-based architecture and represent critical bottlenecks for further performance improvements. To overcome these challenges, our proposed TAPL is a novel WSVAD framework that integrates temporal modeling and cross-modal alignment. First, we introduce a lightweight temporal adapter with multi-head attention to capture long-range spatiotemporal dependencies in videos. Second, we design a dual-branch text prompt mechanism that combines static expert prompt with dynamic learnable prompt, enhancing cross-modal alignment between video and text features. Extensive experiments on UCF-Crime and XD-Violence benchmarks validate our approach’s effectiveness. Our method achieves state-of-the-art performance with 88% AUC on UCF-Crime and 85.26% AP on XD-Violence, outperforming prior work. This work offers a robust solution for weakly-supervised video anomaly detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Temporal-aware prompt learning for weakly-supervised video anomaly detection

  • Lin Yuan,
  • Xun Duan,
  • Guangqian Kong,
  • Huiyun Long

摘要

Weakly-Supervised Video Anomaly Detection (WSVAD) is a critical technology in intelligent surveillance, enabling precise temporal localization of anomalies using only video-level labels. While CLIP-based approaches have significantly advanced the field, two key limitations persist: (1) inadequate temporal modeling capabilities, and (2) suboptimal utilization of textual information. These constraints stem fundamentally from CLIP’s static image-based architecture and represent critical bottlenecks for further performance improvements. To overcome these challenges, our proposed TAPL is a novel WSVAD framework that integrates temporal modeling and cross-modal alignment. First, we introduce a lightweight temporal adapter with multi-head attention to capture long-range spatiotemporal dependencies in videos. Second, we design a dual-branch text prompt mechanism that combines static expert prompt with dynamic learnable prompt, enhancing cross-modal alignment between video and text features. Extensive experiments on UCF-Crime and XD-Violence benchmarks validate our approach’s effectiveness. Our method achieves state-of-the-art performance with 88% AUC on UCF-Crime and 85.26% AP on XD-Violence, outperforming prior work. This work offers a robust solution for weakly-supervised video anomaly detection.