Accident-CLIP: Text-Video Benchmarking for Fine-Grained Accident Classification in Driving Scenes
摘要
Road accident classification which is an essential but rarely explored problem in the safe driving field. The core issue of accident classification is to learn the feature representation for partitioning different kinds of accidents. Compared with accident detection or anticipation only with occurrence probability, feature representation learning in accident classification is more challenging because of the extremely imbalanced accident categories. In addition, the severe light or weather conditions, various occasions, and complex crashing-object movement exacerbate the challenges. In this work, we form a text-video benchmark for fine-grained road accident classification (named Accident-CLIP). Accident-CLIP owns 13,669 dashcam videos with 58 kinds of accidents, where each accident is annotated with the text description of the accident type and the accident window. In the benchmarking stage, six state-of-the-art methods are evaluated from different video frame sampling methods, frame mixup strategy, input frame length, and the adaptation for long-tailed accident distribution. From experiments, we observe that current video classification models need a large space (the best Top-1 value is 41.39% on 2000 testing videos) to adapt to the extremely imbalanced road accident classification, and the formed Accident-CLIP benchmark provides a promising evaluation platform.