Improving Anomaly Scene Recognition with Large Vision-Language Models
摘要
Vision-based anomaly scene recognition is important for plenty of applications such as surveillance and security. An efficient way to achieve anomaly scene recognition is to use image-based methods. However, the accuracy achieved by existing image-based methods is relatively low. In this work, we introduce AnoSRe, a novel method designed to accurately recognize anomaly scenes from images. AnoSRe takes advantage of the pre-trained large vision-language model to improve the accuracy of popular deep learning networks for anomaly scene recognition. Experiments conducted on two datasets show that the proposed method outperforms state-of-the-art methods, and improves the accuracy by from 5% to 23% on UCF-Crime-Image and IDSR datasets.