Temporal-aware prompt learning for weakly-supervised video anomaly detection
摘要
Weakly-Supervised Video Anomaly Detection (WSVAD) is a critical technology in intelligent surveillance, enabling precise temporal localization of anomalies using only video-level labels. While CLIP-based approaches have significantly advanced the field, two key limitations persist: (1) inadequate temporal modeling capabilities, and (2) suboptimal utilization of textual information. These constraints stem fundamentally from CLIP’s static image-based architecture and represent critical bottlenecks for further performance improvements. To overcome these challenges, our proposed TAPL is a novel WSVAD framework that integrates temporal modeling and cross-modal alignment. First, we introduce a lightweight temporal adapter with multi-head attention to capture long-range spatiotemporal dependencies in videos. Second, we design a dual-branch text prompt mechanism that combines static expert prompt with dynamic learnable prompt, enhancing cross-modal alignment between video and text features. Extensive experiments on UCF-Crime and XD-Violence benchmarks validate our approach’s effectiveness. Our method achieves state-of-the-art performance with 88% AUC on UCF-Crime and 85.26% AP on XD-Violence, outperforming prior work. This work offers a robust solution for weakly-supervised video anomaly detection.