Self-supervised Traffic Accident Detection by Motion-Conditioned Diffusive Frame Prediction
摘要
The detection of anomalies using dashcam videos is crucial for autonomous driving or driver assistance systems. The scarcity of diverse accident videos and the complex environment variations during driving significant struggle for accident detection. Leveraging the advancements in frame prediction-based accident detection methods and diffusion models, we propose a frame prediction framework based on motion-conditioned diffusion (MCD-TAD). This framework combines optical flow features with appearance features in consecutive video frames using a latent diffusion model to better capture and utilize spatiotemporal cues. Extensive evaluations on two large-scale accident datasets, namely AnAn Accident Detection (A3D) dataset and DADA-2000 dataset, validate the effectiveness of the MCD-TAD for traffic accident detection in dashcam videos.