ViCap-AD: video caption-based weakly supervised video anomaly detection
摘要
Anomaly detection becomes increasingly critical amidst rising crime rates and concerns for public safety. Traditional unsupervised video anomaly detection methods primarily focused on normal data, limiting their ability to achieve optimal performance due to their inability to effectively utilize abnormal data. Weakly supervised video anomaly detection methods addressed some of these limitations but still struggled to leverage anomalous video labels effectively, often being susceptible to noise in classification scores. In this paper, we propose a novel approach named Video Caption Anomaly Detector (ViCap-AD), which leverages video captions alongside a combination of BERT and the multiple instance learning (MIL) framework for anomaly detection. ViCap-AD integrates video captions generated using CLIP4Clip with video features within the MIL framework augmented by BERT. In our experimental evaluations on the UCF-Crime and XD-Violence datasets, ViCap-AD achieves state-of-the-art performance, achieving AUC scores of 87.20% and 85.02%, respectively. These results underscore the robustness and effectiveness of our approach across different datasets, demonstrating its powerful performance and stability. This paper contributes a significant advancement in anomaly detection methodologies, highlighting the potential of ViCap-AD to enhance anomaly detection accuracy and reliability in real-world applications.