A novel lightweight dual-stream recurrent transformer model for anomaly detection in driving dashcam videos
摘要
In autonomous vehicles, it is significant to understand the surrounding environmental context to ensure reliability and safety. To achieve this, detecting a wide range of anomalies surrounding vehicles is crucial. This paper proposes a new robust and efficient model, called lightweight dual-stream recurrent transformer (LDSRT), for anomaly detection in driving dashcam videos. This model competes with state-of-the-art models, tackling their limited generalization problem, and providing reliable timely detection. The new model introduces an enhanced dual-stream network combining an improved MobileV3-Small network and an ameliorated convolution-3d network. This dual-stream design efficiently produces a compact representation of detailed static visual-spatial, temporal, and dynamic motion features. In addition, the model integrates a recurrent-enhanced transformer network for effectively understanding and learning the contextual relationships and temporal dependencies across driving video frames. This assures a comprehensive analysis of appearance and motion cues, both temporally and contextually. The model is then trained on a new integrated and diverse large dataset: DoTA-StreetAccident. The DoTA-StreetAccident dataset depends on merging two datasets: Detection of Traffic Anomaly (DoTA) and StreetAccident. Further, the LDSRT model’s effectiveness is extensively evaluated on the DoTA-StreetAccident and RetoTruck datasets, confirming its robustness across various driving anomaly detection tasks. Several evaluation metrics are utilized, and the analysis of experimental results and computational cost demonstrates that the proposed model significantly outperforms state-of-the-art models.