EM-OFRP: enhanced memory-based optical flow reconstruction and variational prediction for video anomaly detection
摘要
Video anomaly detection is a primary task in video surveillance systems, with widespread applications in maintaining social security, industrial production safety, and other fields. However, existing video anomaly detection models face significant challenges in accurately capturing fine-grained key features and reducing feature information loss during encoding and decoding processes. This paper proposes an enhanced memory method that integrates a Transformer-based convolutional attention mechanism with a multi-level memory-enhanced network. This approach effectively enhances the model’s ability to capture subtle changes and long-term dependencies, strengthening its perception of contextual information, thereby achieving accurate extraction of key features and preservation of detailed information. Furthermore, through a feature enhancement bridging module using skip connections, features extracted in the encoder are non-linearly processed and transferred to the decoder, thereby compensating for the semantic gap between encoder and decoder features. This design not only accelerates the flow of information, but also enables the decoder to fully exploit the rich memory content of the encoder when generating prediction results, further improving the accuracy and detail fidelity of the predictions. Finally, the hybrid model proposed in this paper further enhances the differences between anomalies and normal events through the reconstruction module, facilitating the detection of anomalies in videos. The experimental results show that the average frame-level AUC of the proposed model reaches 98.5, 91.2 and 77.8% on the benchmark video anomaly detection datasets UCSD Ped2, CUHK Avenue, and ShanghaiTech, fully validating the effectiveness and superiority of the model.